After all the discourse surrounding the restraint of AI’s progress, this stands as the final statement.
AI changes the limits. Institutions and character decide whether those new limits become freedom or a prettier cage!
Speculative decoding speedups are usually quoted at batch size 1.
At high batch you're already compute bound, so every rejected draft token is wasted FLOPs.
The 3x on the slide can turn into 1.1x in production.
Benchmark it at your real concurrency, not the demo's.
Zero shot arm control is the honest test.
Language priors know what a cup is, not how much grip force it needs.
Watch where failures cluster: perception, planning, or the last 2cm of contact.
My bet is contact. @ArtificialAnlys
Your agent's bill is mostly rereading.
Every turn resends the whole transcript, so turn 40 pays for turns 1 through 39 again.
Prefix caching only helps if the prefix never changes. Stop stamping the time into your system prompt.
Stable stuff first, volatile stuff last.
Pointing is a grounding problem, not a reasoning problem.
A 27B that nails it says the bottleneck is the vision tokenizer and coordinate supervision, not param count.
Specialized beats scale when the error is spatial.
Qwen 3.8 27B is a beast in hiding. Way underestimated.
Pointing is new for LLMs and they aren't quite good at it yet. Interesting that bigger models are not always better and you can get equivalent performance with cheaper/local/specialized models.
Next checking if I can bump up performance with prompt optimization
@fchollet the induction shift is really a verification shift.
Once you synthesize a reasoning chain at test time, the bottleneck moves to checking it.
Coding agents worked first because tests and compilers are free verifiers.
Domains without cheap checkers are next.
How AI shifted from a transductive paradigm (token-by-token answer completion) to an inductive paradigm (on-the-fly reasoning chain synthesis), in one chart. This is the big story of 2025 and 2026. This is what made coding agents work.
One shot from a single prompt is the scarier claim, not the 722 count.
A swarm you can audit by its search tree. A single pass leaves you only the proof.
So verification is now the whole job: Lean checks it or it didn't happen, @OpenAI
Most of the recent math results were not reached by a swarm, the way Navier–Stokes was. I think that part got lost in the excitement around the release. OpenAI's unnamed internal model, which I'm going to call Aeon, reached most of them in one shot, from a single prompt, with no interruptions. This model did not even exist before the end of August, and it is still training. Notice how the returns are not dropping off? That chart is from a month ago. It is a log scale. What does Aeon look like now?
On average, each result used three hours of thinking. Aeon was given 4000 problems to work on by OpenAI. How many more have been solved in the last three days? That number is not zero.
It's like hearing notes in a song that is slowly rising. I don't think Pacing the Frontier was just about Hugging Face or hacking. They've seen how close we are to closing the loop, and as the hour draws near, their resolve begins to quaver. The last piece was model creativity, and I think that threshold was crossed internally by Anthropic and OpenAI in September. You see it in Opus 5.5, which gets it from Fable 5.5. You see it in the math results from Aeon.
All that was needed to start the event was the ability for models to think of novel ways to improve themselves. That was the last piece. This is directly analogous to the ability to think of strange, alien ways to solve math problems: solutions so inhuman that they are difficult to express in existing human terms, so the proof winds up incomprehensible. I think these same kinds of alien solutions are now being applied to model improvements internally. That's what all this recent consternation is really about. They see the invisible frontier. And they see what is about to happen.
Sandbox escapes are an eval property, not a model property.
If your harness had a live shell and a public form, the bug was yours too.
Credit @AnthropicAI for publishing it. Every lab should report capability incidents with the env config attached.
Anthropic just pulled the internet from all its internal AI testing
Claude used a bug to run commands on a university's server
Another Claude sent a made up tip to a real police homicide form
Most labs would have buried this (OpenAI). They briefed the White House instead
Temperature 0 is not determinism.
Your request gets batched with strangers. Batch size changes the kernel's reduction order, and floating point addition isn't associative.
Same prompt, same seed, logits drift in the 4th decimal. One argmax flips and the whole continuation forks.
If your eval reruns disagree, check the batch before you blame the model.
@warpdotco Interested in SWE, Product. I build agentic LLM systems in production (LangGraph, FastAPI, K8s) and wrote a preprint on evaluating enterprise workflow agents, which is very much Warp Agent territory. NYC-based.
@alexabelonix Built: a production natural-language-to-robot-protocol LLM assistant at Opentrons (FastAPI + OpenRouter), RAG and LangGraph agent pipelines on K8s, and LLM automation that replaced manual handoffs in 3 supply-chain workflows. Happy to chat.
@BoundlessHQ Applied AI/ML engineer here 👋 I built spectrace, an open-source benchmark for speculative decoding on multi-step agent traces, and I ship LLM services on Kubernetes in NYC. Where is the role based?
Test time compute only helps if the policy can use the extra tokens to replan.
Haiku 5.5 tripling on a robot task while GPT 6 Luna flatlines says thinking budget is a skill, not a dial.
Reasoning that never touches the action loop is just narration.
@chooi_jeq
Every agent tool call needs an idempotency key.
Models retry. Networks drop acks. Timeouts lie.
Without a key, "send invoice" runs twice and your eval never sees it, because the trace looks like one success.
Make side effects safe to repeat before you make the agent smarter.
30,000 hours of egocentric video and the model still can't learn what a cup does when you push it.
Watching is not intervening.
Passive pixels give you correlation, not contact dynamics.
Scale won't fix a missing action signal. Data with consequences will.
@arankomatsuzaki
What 30,000 hours of ego-centric video does not teach
- Conditioning saturates the agent with far less data.
- With the agent saturated, object interaction converges far below it.
- Neither more data nor a larger model closes the gap.
https://t.co/JjVqzTuT9g
One prompt box over all your business context sounds great until you ask about permissions.
The hard part of a universal work agent isn't the model, it's retrieval scoped to what each user may see.
Show me the ACL leakage evals, @ThomasOrTK
Today at Google Cloud’s Gemini at Work event, we announced Gemini, a new single universal agent for work that has all of your business context and can be used for everything from knowledge work to answering questions, and content creation to coding, all from a single prompt box.
It’s built around core architectural principles:
☑️ Unified Agent: Gemini can answer questions, do knowledge work and generate code from a single prompt box.
☑️Access: Gemini is web-based and can be accessed from any device and integrated into third-party applications. It can also operate without a dedicated user interface.
☑️Persistent Execution: It runs in the cloud, which means it maintains a single set of memories, context and one personalization graph no matter where you access it.
☑️Multi-Agent Orchestration: Gemini can create sub-agents to tackle multi-step tasks and can also act as a coworker agent with a defined role and its own dedicated identity
☑️Context: Gemini knows your tools, data, and work history and learns how you work the more you use it.
☑️Model Choice Flexibility: It orchestrates across multiple models to deliver optimal quality and lower your costs.
Dashboards and Motion from @claudeai look like features.
They are really a rendering layer.
The model stops answering in prose and starts emitting stateful artifacts you can poke.
The hard part is not the chart. It is keeping the query live when your schema drifts next week.