@AmenaiSabuwala@claudeai everyone's asking skip what, but i read it as trying @claudeai without an account first, @ChatGPT already lets you do that, so it's a bigger change than the background tbh. @AmenaiSabuwala was that the idea?
@wuweiweiwu 2 years making tests easier to write and then launching the thing that means you never touch them again. respect for eating your own product lol
@ParallelAILabs Thanks. The case I'm still working out is an agent editing on someone's behalf, is that a square, a circle, or both? Feels like "who did it" and "who asked for it" both matter there
The most interesting part of the GPT-6 Sol launch isn't the price cut.
It's one reply, shown before and after. The old one says it's done. The new one says what it changed and what it checked: desktop, narrow mobile, the back button.
OpenAI called that the better answer.
Agents are learning what good products already knew. Telling people what you checked is part of the job.
OpenAI: APIs and tools for building AI products by @sama https://t.co/8Pvls95ntu
When people and AI agents work on the same doc, "who did this?" can't be a guess.
Been designing a brand where the answer is a shape. People are circles. Agents are rounded squares. The mark is both, side by side, never merged.
Full concept on Behance soon
Three models tied at 97%.
So the model isn't the difference anymore. The 3% is. What your product does when the answer's wrong is the part users actually remember, and it's the one part you can't buy from a lab
@ClaudeDevs Putting cost per task in /usage instead of a pricing page is the right call, I think. One layer I'd add for plan users: how much of the 5h window a task eats. A lot of the replies here think in windows, not dollars.
@ycombinator What doesn't carry over from coding agents is undo. A bad code step is a revert, a bad robot step is something on the floor. So I think the approval moment has to move from after the action to before it
An AI spent 20 minutes thinking about a pelican on a bicycle this week, then said nothing.
@simonw gives every new model the same test. On Tuesday he gave it to Claude Opus 5.5 on its highest thinking setting. It planned the beak, checked the pelican's shin length, planned a fish in the basket, then ran out of room before drawing anything. Twice. Nearly 20 minutes and $2.56 each time.
Same day, the price war started. Opus 5.5 got 20% cheaper with cached input down 60%, and OpenAI shipped GPT-6 Sol and Luna at half the price of what they replace. Long agent runs just got a lot cheaper.
So I think products are about to run longer, and more people will sit through their own version of those 20 minutes.
That's the screen most products don't design. What's it doing, how far along is it, what's it costing me, can I stop it, what do I get if it fails?
Simon got a blog post out of his 20 minutes. Your users probably just get a spinner
@onethomaschan@ycombinator #1 has a quiet bonus i think. the launch post ends up being the clearest pitch you'll write, because you wrote it before polishing. worth watching which lines people quote back and moving those onto the site