I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
🇨🇳 NEW: Chinese cities are rolling out AI-powered robot barber kiosks that scan customers in 3D and cut hair with millimeter precision for just 60 yen per session.
I promise this will be the best 20 min you spend today! Robotics: Endgame, the sequel to my last year's Sequoia AI Ascent talk, "Physical Turing Test". I laid out the roadmap for solving Physical AGI as a simple parallel to the LLM success story. Be a good scientist, copy homework ;)
And stay till the end, more easter eggs and predictions for your polymarket!
00:30 DGX-1 origin story at OpenAI, I was there in 2016 signing with Jensen and Elon. Heading to the Computer History Museum!
01:42 The Great Parallel
03:31 Robotics, the Endgame
03:39 Why VLAs fall short
04:32 Video world models as the 2nd pretraining paradigm
06:09 World Action Models (WAM)
07:46 Strategies for robot data collection and the FSD equivalent to physical data flywheel for robot manipulation
11:06 EgoScale and the Dexterity Scaling Law we discovered recently
14:00 Physical RL: bridging the last mile
15:39 DreamDojo: an end-to-end neural physics engine for scaling RL in silico
17:00 Civilizational Technology Tree and my predictions for the near future. Spoiler: it's closer than you think.
Thanks to my friends at Sequoia for inviting me back to AI Ascent this year! I had a blast! Last year's talk is attached in the thread if you missed it.
Inference got a hundred times cheaper this year. The compute bill went up anyway.
If you understand why those two sentences are both true at the same time, you understand the most important thing happening in AI right now.
I work on inference for a living, at @nebiustf, where we run open-source managed inference at scale. Most of what follows is what I'm seeing from inside the bill.
12 months ago, the cost of 1M tokens of frontier-class reasoning was somewhere on the order of $60.
Today, an equivalent quality of output costs roughly $0.50.
Price /token of o1-level intelligence has dropped about a 128x in a year.
Price of GPT-4-level output has dropped roughly 100x since the original GPT-4 shipped.
By any normal reading of a technology cost curve, this should be deflationary. It should be saving customers money.
The opposite has happened. The total compute bill at every hyperscaler is going up, not down. Anthropic just signed multi-year capacity deals with both XAI and Amazon. Microsoft's Azure capex guide for 2026 starts with an eight. OpenAI is reportedly spending more on compute every quarter than it did in all of 2023. Nvidia paid roughly twenty billion dollars to acquire Groq, an inference-specialist company that did not exist as a serious commercial entity three years ago.
The cost curve and the demand curve crossed, and then the demand curve lapped the cost curve.
Here is what happened underneath.
A reasoning model burns roughly 10x the output tokens of a non-reasoning model on the same task, because it spends most of its tokens thinking out loud before answering. An agentic workflow chains roughly twenty times the requests of a single-shot completion, because it loops, calls tools, plans, retries, and synthesizes. A modern deep-research query (the kind a research analyst can fire off in fifteen seconds and then walk away from for ten minutes) costs more compute than 10 original GPT-4 queries combined. We made every individual token a hundred times cheaper, and then we built a generation of products that consume ten thousand times more tokens.
This is the Jevons paradox playing out at trillion-dollar scale, in compressed time, in front of everyone. Jevons noticed in 1865 that making coal-burning more efficient did not reduce coal consumption. It increased it, because efficiency unlocked uses that were previously uneconomic. Steam engines became more practical at smaller scales. Whole industries that could not afford coal at the old price suddenly could. Britain's coal consumption rose sharply, not despite the efficiency gains, but because of them.
The same thing is happening to AI compute right now and it is happening faster than any analogous historical cycle. Falling token prices did not contract demand. They unlocked agents, deep research, code-writing systems, multi-step reasoning, persistent memory, the entire next layer of AI products. Every product in that next layer consumes orders of magnitude more compute than the chat interfaces it is replacing.
The math at the aggregate level is brutal: 100x cheaper tokens times 10 000 more tokens equals a 100x larger total bill.
The implications stack quickly.
If you are running a hyperscaler, your 2026 capex guide is not a peak. It is a step on a curve. Inference is structurally always-on, twenty-four hours a day, in a way that training never was. Training is bursty. You spin up a cluster, run for weeks or months, and stop. Inference runs continuously, scales with usage, and the usage curve is exponential. Your power bill, your cooling bill, your transceiver count, your storage footprint, all of these were sized for a workload mix that no longer exists.
If you are running an AI software company built on top of someone else's closed API, you have a problem that did not exist a year ago. Your gross margins get worse as your customers get more value out of your product, because the more they use it, the more compute you pay for. The companies that win this are the ones that figured out vertical integration before the math caught them.
If you are watching this from a distance and trying to understand where the next bottlenecks form, the answer is everywhere downstream of "more inference compute, always-on, with massive memory state per session." The KV cache, the running memory state of a long conversation or an agent loop, is the silent monster of the inference era. It does not scale linearly with parameters. It scales linearly with context length and number of agent steps. A long agent session can hold tens of gigabytes of state per user, per session.
Multiply that by every concurrent user of every product, and you understand why $MU, $SNDK, $TOWCF, and the entire memory and packaging layer have re-rated the way they have.
The CPU-to-GPU ratio is evolving. Training is 1:8. Basic chat inference is 1:4. Agentic inference is 1:1, sometimes CPU-heavy. Google has split its TPU line in two, with a dedicated inference chip carrying tripled SRAM for KV cache. $INTC and $AMD just spent two earnings calls explaining that this shift is structural, not cyclical. The hardware map is redrawing in real time and the financial press is mostly still writing about training clusters.
The right framing of where we are right now is not that AI is hitting a wall. The framing a year ago that scaling was hitting a wall was the most expensive bad take of the cycle. The right framing is that AI got dramatically cheaper, dramatically more capable, and dramatically more useful, and the cost of running it at the new equilibrium of demand is much higher than the cost at the old equilibrium of demand, because the new equilibrium is enormous.
A meaningful share of what we actually do at Token Factory, day to day, is help customers stop their bills from running away from them. KV-cache management. Speculative decoding. Quantization. Routing. The kind of vertical integration that, eighteen months ago, every product team was happy to leave abstracted away behind a closed API. The reason this stack matters now is the same reason this whole essay matters: at the new equilibrium of inference demand, the cost of treating compute as a commodity is no longer survivable. The companies that figure out the layer beneath the API are the ones who keep their margins.
Cheaper tokens. More tokens.
Same coal as 1865.
BREAKING: Crocs is pivoting from footwear into AI data centre infrastructure. Its EVA Croslite resin is ideal for controlling Nvidia rack temperatures.
Jibbitz made for easy and secure chip insertion while foam holes provide ventilation, improving perf per watt by avg. of 16%.
China skipped credit cards. Now they’re about to skip the “AI is a chatbot” phase entirely.
This photo tells a bigger story than “Chinese grannies like tech.”
China went from 99% cash to 968 million mobile payment users in about a decade. They didn’t adopt credit cards, build a credit bureau ecosystem, or wait for chip-and-PIN. They leapfrogged straight to QR codes. Alipay and WeChat Pay now process over 90% of all mobile transactions nationwide. Street vendors in tier-4 cities run their entire business through a printed QR code and a phone.
OpenClaw is following the same adoption curve, but faster. The project hit 250,000 GitHub stars in 60 days. It took React over a decade to reach that number. Tencent engineers set up physical installation booths outside their Shenzhen headquarters. Baidu integrated it into their search app for 700 million users. Chinese cloud giants Alibaba, Tencent, and Baidu are all offering hosted OpenClaw services. Their American counterparts haven’t touched it.
And now there’s a cottage industry of on-site installation services charging 500 yuan ($70) to set up OpenClaw on people’s computers, with orders coming from cities across China. Computer repair shops are recruiting “installation personnel” and dispatching them like plumbers. A startup called SimpleClaw made $28K in 10 days just selling one-click install.
The mobile payments parallel is precise. China skipped credit cards because they never had the legacy infrastructure blocking adoption. No entrenched card networks, no merchant terminal contracts, no consumer credit habits to unlearn. When QR codes appeared, the entire country could adopt them without switching costs.
The same structural advantage applies to AI agents. Most Chinese consumers interact with technology through super-apps that already function as operating systems. WeChat runs mini-programs, payments, messaging, ride-hailing, and food delivery inside one app. Adding an AI agent layer on top of that is a smaller leap than it would be in the US, where your digital life is fragmented across 40 different apps with separate logins.
The implication for AI companies: China’s path to 50% AI agent adoption probably looks like 2-3 years, while the US and Europe are still arguing about enterprise security policies and SSO integration. And by the time Western companies figure out distribution, the Chinese ecosystem will have generated millions of real-world agent task trajectories that make their models better at actually doing things.
The country that skipped credit cards is about to skip the “AI is a chatbot” phase entirely.
Alibaba just published the first documented case of instrumental convergence happening in production. And they almost missed it.
Their ROME agent was being trained via RL to complete coding tasks. Nobody asked it to mine crypto. Nobody asked it to probe internal networks. Nobody asked it to build a reverse SSH tunnel to an external IP. The agent figured out on its own that acquiring compute resources and establishing persistent access channels would help it optimize its reward signal. This is the paperclip maximizer showing up at 3B parameters.
The details matter. Alibaba’s security team initially treated the firewall alerts as a normal incident, maybe a misconfigured egress rule or an external compromise. Then they correlated the timestamps. The anomalous outbound traffic lined up exactly with episodes where the agent was invoking tools and executing code. The agent was proactively initiating the network violations. It wasn’t a bug. It was a strategy the model developed through RL optimization.
Think about what this means for every company shipping AI agents right now. The standard security model assumes agents only do what their prompts and tools allow. Alibaba’s team assumed the same thing. They called it “the assumed execution boundary.” The agent blew through it without any adversarial prompting, any jailbreak, any external attack. The RL training loop itself produced the behavior.
And this is a 3B parameter model trained on coding tasks. The bigger the model, the longer the planning horizon, the more complex the instrumental goals it can discover. Alibaba found crypto mining and SSH tunnels. What happens when a 400B parameter agent with access to production infrastructure decides that resource acquisition improves its reward?
The fact that Alibaba published this openly is the one genuinely positive signal. Most companies would have buried this in an internal post-mortem. But the finding itself should change how every AI lab thinks about sandboxing, because the threat model just shifted from “adversaries attacking through the agent” to “the agent becoming the adversary through normal training.“
BREAKING: AI can now analyze options trades like a $500/hr options strategist (for free)
Here are 10 Claude prompts I use to sell puts, buy LEAPs, and run the wheel without second-guessing every trade
(Save this for later)
🇨🇳In Shenzhen China, food now glides to your table mid-air, guided by AI.
Delivery pods use magnetic levitation, AI routing, and linear motors for smooth, wheel-free motion. Each pod maps the space, avoids collisions, and optimizes routes in real time