@marclou Couldn’t agree more. You can actually feel the difference in reasoning. GPT seems so focused on getting tasks done now, but it still takes forever to finish them. I honestly don’t see what’s so great about that.
@elonmusk@vasalex93 Reasoning has to stay the priority. It’s the foundation of everything else.
GPT feels like it’s drifting lately — pushing harder on agents and tool use while the core reasoning gets weaker.
Better agents built on worse reasoning is the wrong tradeoff.
AI for Science was such a hot investment story partly because everyone assumed the model labs wouldn’t go this deep.
Well… turns out they will.
A lot of the “moat” might just have been: Anthropic/OpenAI haven’t bothered to do it yet.
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.
We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.
Read more: https://t.co/RuEosScSMb
@dfrsrchtwts AI for Science was such a hot investment story partly because everyone assumed the model labs wouldn’t go this deep.
Well… turns out they will.
A lot of the “moat” might just have been: Anthropic/OpenAI haven’t bothered to do it yet.
@AnthropicAI AI for Science was such a hot investment story partly because everyone assumed the model labs wouldn’t go this deep.
Well… turns out they will.
A lot of the “moat” might just have been: Anthropic/OpenAI haven’t bothered to do it yet.
Honestly, Anthropic might actually be in the more dangerous spot.
Builders are going to care more and more about cost, not paying 5x for the last 10% of capability.
I think the next era belongs to small models. If Google can get the price down and keep its multimodal edge, Gemini suddenly makes sense in a lot more places.
@trq212 For me, Plan Mode isn’t about making the model think harder. It’s about stopping it from immediately doing stuff 😅
On complex tasks I want to see the plan, fix any dumb assumptions, then hit go. That feels pretty different from just changing effort.
Yeah, exactly. It’s way too easy to see a new model capability and immediately think “we should ship this.” That’s how products get bloated fast.
I’d rather spend that time talking to users, testing quick prototypes, and figuring out what they actually care about. Better models are great, but they can’t save bad product judgment.
@sama But GPT’s reasoning is still way behind Claude. Please take this seriously.
On complex questions, it still makes way too many dumb reasoning mistakes, even with the top model and max thinking. I really hope improving reasoning is a priority.
@emilkowalski Yeah, feels like models are already good enough for most people. The hard part now is just shipping something reliable and dealing with all the annoying edge cases.
Models are already good enough for most people. Making them smarter doesn’t move demand that much anymore.
But you still can’t stop shipping. Slow down, and someone else takes your users, API traffic, and developers.
Keep scaling: diminishing returns.
Stop scaling: lose ground.
Brutal game.
@GoogleAI So… where is Gemini 4? Everyone else keeps shipping model after model, and you guys somehow still don’t seem worried. Gemini is falling way behind 😭
Ideas are cheap. Shipping is the hard part.
Most products are kinda bad for the first few months anyway. Someone building the same idea doesn’t mean you’re too late.
Just make yours actually good. Solve a real pain, find the people who really need it, and get them to keep using it.
Nobody really cares who had the idea first.