We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
Every startup should have a daily markdown file called "what_the_market_is_telling_us.md"
It updates every morning from the places where customer truth already lives:
1. Stripe for who pays, upgrades, downgrades, and churns
2. PostHog for what people actually do in the product
3. Intercom or Plain for support tickets/complaints
4. Granola or Gmeet transcriptions for sales calls/ customer interviews
5. HubSpot or Salesforce for CRM notes/lost deal reasons
6. Linear, Jira, or GitHub Issues for bugs and feature requests etc
7. Ideabrowser MCP for outside market signal: startup ideas, trend reports, social/search demand, AI research reports, and builder prompts that show what people are starting to want before it shows up in your own customer data.
Basically, the file should notice what changed in the business this week and not just be this summary of here’s what happened (which I think a lot of people have their agents do).
Why this is valuable:
1. Maybe new buyers are using different words than they were a month ago.
2. Maybe trial users are getting stuck in the same place.
3. Maybe upgraded customers all touched one feature right before they paid.
4. Maybe churned customers keep mentioning setup confusion.
5. Maybe sales calls are suddenly losing to a competitor you used to beat.
6. Maybe support tickets are revealing a workflow your product accidentally became responsible for.
You get the point.
The fastest way to PMF is understanding customers better than anyone else, and the highest signal customer insight is usually a change in behavior.
So I’d have the agent update the file every morning with the pattern it found, the receipts behind it, and the product or GTM decision it might affect.
For example:
“3 customers who churned this week all mentioned setup confusion, and 2 of them never invited a teammate.
This looks more like an activation problem than a pricing problem, so I’d look at team invite and onboarding before building another analytics feature.”
A little helpful tip for all those out there looking to get more from their LLMs.
See what needs you next.
The new Activity view in the ChatGPT desktop app brings together conversations that need your attention and recent updates across your projects.
Hosting our own Buzz w/ Cloud Agents
(Buzz Video #2:)
For the past week, @vishal_dubey and I have been using Buzz.
Yesterday, Vishal used Fable 5 (inside Buzz) to create our own hosted version of Buzz with built-in cloud agents that never die.
These agents can spin up new agents that all share the same skills, API keys, workspace, files, automations, and more.
This is an EXPERIMENT. We’re trying to find the best setup for businesses to create teams of agents and humans that work together.
Part 1: What we like and don’t like about Buzz
Part 2: Vishal’s hosted version of Buzz
⭐ Part 1
00:00 Intro
01:06 What Is a Harness?
02:49 What We Like About Buzz
04:17 Agents Collaborating Together
06:29 Custom Agents and Personas
08:13 Spinning Up Agents and Channels
10:34 What We Don't Like about Buzz
⭐ Part 2
14:03 Vishals Version of Buzz with Cloud Agents
19:39 Shared Skills and Connections
23:18 Setting Up a Daily Research Automation
26:08 Buzz Mobile App
30:53 Treat the agents like their own employee
35:28 Using Any Model (Fable, 5.6, Kimi)
37:25 What Vishal wish Buzz Had
Opus 5 is getting eerily good at creating product videos with camera moves like zooms, perspective shifts, and seamless transitions.
This 8-minute video breaks down the workflow:
https://t.co/aI9QiRAWwE
🆕 First Steps Toward Automated AI Research
https://t.co/ypWRfqy65k
Humanity advances by trying things, finding the shortcomings, and fixing them. @RichardSocher keynotes our first-ever Autoresearch track to show how @Recursive_SI is building a Eureka machine for recursive self improvement... and how far we have yet to go.
i just found a great way to AI-generate UGC videos for your product automatically via. claude code...
first find a video that you wanna take inspiration from (make sure same style is applicable to your own product)
ask claude code to connect to higgsfield CLI & ask the agent to use the built-in video analyzer to analyze the video
ask it to recreate it with seedance/omni (make sure you add in your own product context too, the more context the more accurate the video will be)
you can edit in clips of your own software/app to make it cohesive & realistic
i'm fully sold that all marketing workflows will be done through the terminal, it's happening already
I’m sick of reading AI slop, so today I’m open-sourcing my /no-ai-slop skill that removes 20+ slop patterns from any piece of writing.
📌 Get the free skill here: https://t.co/YshUga82ks
If you find it useful, please consider starring the repo so more people can find it.
Why I built the skill:
I use AI to edit my writing because it helps me fix spelling, grammar, and clarity.
But even the best models keep producing the same slop that this skill removes:
→ Binary contrasts: “It’s not X. It’s Y.”
→ Throat-clearing openers: “Here’s what nobody tells you.”
→ Fake-profound endings: “The future isn’t coming. It’s already here.”
Use this skill responsibly.
I always write a first draft manually before iterating with AI on edits and I make sure to do another manual pass at the end as well.
That’s in contrast to using AI to automate pumping out slop end-to-end.
📌 Read my full post for more on how I try to use AI responsibly to edit without giving into the dark side: https://t.co/C68RSabtb0
chatgpt work is remarkable, and "work" undersells it.
from my phone i sent:
"use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready."
it...just worked.
I tried @jack's Buzz.
It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on.
The video below shows how it works, and some of my thoughts on the process and platform, e.g.:
- Create and interact with agents on top of any harness (claude code, codex, pi, etc.)
- Choose which models agents use, including local ones
- Agents can delegate work and work in parallel in git worktrees
- Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history)
- You can share AI compute within a community
- It's completely open-source and decentralized
Things I like:
- Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys.
- Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it.
- It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future.
- It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool.
Things I didn't like:
- You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great.
- It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead.
Verdict:
- I really like it so far and can genuinely imagine working with a team this way.
- It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect.
- The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
@josh_bickett Haha thanks Josh!
https://t.co/YxEHek9L8p for anyone who wants to try the workflow... it's the most powerful way to leverage AI.
I built it for myself, and it's not really a product, so expect rough edges (I'm happy to help if anyone gets stuck though, just tweet at me).
Matt Shumer (@mattshumer_) really is living in the future.
He just showed me his setup. He is running a whole company of agents via workbench .md.
> agents working across laptops and cloud VMs
> one chief of staff checking in on each agent
> all the agents can communicate with each other via workbench
> Matt chats directly with chief of staff to get updates and share input
I've not seen anything like it. I need to up my agent game.
You can now ask Claude about the Anthropic Economic Index, our public dataset measuring how AI is used across the economy.
Ask which occupations use AI the most, or what kinds of tasks people are automating, and the answers draw directly from the Index data.
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.