LLMs aren't bad at coding. They're bad at asking questions.
Fresh results on the IdeaAMBIG benchmark show a massive gap:
โข Spotting an ambiguity on their own: 9.6%
โข Fixing it once pointed out: 80.6%
Why? LLMs are trained to complete text, not to handle uncertainty. When a prompt is vague, they don't pause they guess the most statistically likely intent with 100% confidence. This 71% gap is the single biggest reason autonomous coding agents derail, while human in the loop dev setups feel magic
The next big upgrade for agents won't be better code generation, but better self-doubt
@_KevinTang Old games getting a second life in VR is honestly one of the coolest things happening rn. Astra got one running on my Quest 3 too and I genuinely wasnโt expecting it to work that well
DeepSeek is about to do something pretty funny.
V4.1 Flash apparently beats V4 Pro on capability, speed, cost and total task time.
So until V4.1 Pro ships, they're literally routing Pro requests to Flash and charging Flash prices.
โFlashโ is slowly becoming a pricing tier, not a capability tier
โWhatโs the smartest model?โ is becoming the wrong question.
GPT-6 Astra and Claude Fable 5.1 are tied at 53 on the new Artificial Analysis Intelligence Index.
Yet Astra wins agent automation + terminal work, while Claude wins GDPval, SciCode and HLE.
Same score. Very different brains.
Trying out @spawn right now and Iโm genuinely impressed
You can describe a game build it live in the browser and publish it
Creating is currently free, and games are multiplayer from the start
Really promising !
Recurrent architectures are back in the AI spotlight.
Sapient Intelligence says it has been working on recurrent reasoning since 2025, with HRM and now HRM-Text, an open-source ~1B parameter model.
The wild part: they claim it was trained from scratch on 40B unique tokens for $1,000 in GPU compute.
Theyโre now building PRAXIST to help researchers experiment with new AI architectures.
Recurrence might be having a comeback
Iโm building a Rocket League AI
Itโs currently at 80B steps and watching it improve is honestly so fun.
Would you guys want to see a video of how it plays? ๐