After all the discourse surrounding the restraint of AI’s progress, this stands as the final statement.
AI changes the limits. Institutions and character decide whether those new limits become freedom or a prettier cage!
If AI reaches human level on every single job by training on massive data, it’s still not AGI because it’s far worse than humans at the one thing that matters most: learning.
pass@k is the wrong number for an agent you're about to deploy.
pass@k asks: did any of k tries work?
Production asks: did every one of them work?
If one run succeeds 80% of the time, pass^8 is 0.8^8. About 17%.
Retries can rescue a coding agent. They can't rescue one that already sent the email.
Report pass^k next to pass@1.
1.5T tokens a month sounds insane until you ask how many were cache reads.
An agent loop resends the whole context every step. Most of those tokens are the same prefix read again, billed at roughly a tenth of fresh input.
Multiply raw token count by list price and you get a number that's off by 10x or more.
The real bill is set by cache hit rate. Keep the prefix stable, append only, and the meter barely moves. @theo
I got some issues with this.
The $132m/year is not real at all. Closer to $3m at today's prices, and as low as $1.2k within a year.
Yes, $132m -> $1.2k. Hear me out.
1. The price listed is not accurate at all.
$11m/month is just wrong. Real "api cost" is prolly $100k-$500k per month range.
The 1.5T tokens number includes cache'd tokens. I looked at my own usage, and 33.5B tokens came out to $0.56/m (under 1/10th of Peter's $8/m estimate).
I'm primarily using expensive models like Opus and Astra, and "1.5t tokens" would still have cost me ~$800k. If I used cheaper models and hit cache more, easily could get as low as $100k.
2. Intelligence is getting cheaper incredibly quickly
I know our LLM bills are going up and everyone is stressed about it. Totally reasonable! It's not because intelligence got more expensive.
Artificial Analysis tracks how "cost at a given intelligence level" drops over time. It used to drop at ~10x/year. Now it's at 30x lower in 5 months. (see image)
Assuming the $100k from before, cost could get as low as $3k/month in a few months, and if trends continue, $100/month by this time next year.
Obviously there will be smarter models and we will move to them, but replicating her experience today with equivalent models will be $1.2k/year by next year.
3. Nobody said this is normal
Lauren is the 0.0001% of "agentic engineer". Measuring her work as "$ in -> software out" is reductive and wrong.
Her spend should be looked at as a combination of product/eng, research, FDE work (pstack), and marketing. Worth $350k/day? No. Worth a mil or two a month? Absolutely.
She is exploring what happens when you ignore the cost of intelligence. She's not doing this because Cursor expects every company to spend $10m/year on tokens.
She's doing it because that $10m in tokens is going to cost $10,000 in a year or two.
1.5T tokens a month sounds insane until you ask how many were cache reads.
An agent loop resends the whole context every step. Most of those tokens are the same prefix read again, billed at roughly a tenth of fresh input.
Multiply raw token count by list price and you get a number that's off by 10x or more.
The real bill is set by cache hit rate. Keep the prefix stable, append only, and the meter barely moves. @theo
I got some issues with this.
The $132m/year is not real at all. Closer to $3m at today's prices, and as low as $1.2k within a year.
Yes, $132m -> $1.2k. Hear me out.
1. The price listed is not accurate at all.
$11m/month is just wrong. Real "api cost" is prolly $100k-$500k per month range.
The 1.5T tokens number includes cache'd tokens. I looked at my own usage, and 33.5B tokens came out to $0.56/m (under 1/10th of Peter's $8/m estimate).
I'm primarily using expensive models like Opus and Astra, and "1.5t tokens" would still have cost me ~$800k. If I used cheaper models and hit cache more, easily could get as low as $100k.
2. Intelligence is getting cheaper incredibly quickly
I know our LLM bills are going up and everyone is stressed about it. Totally reasonable! It's not because intelligence got more expensive.
Artificial Analysis tracks how "cost at a given intelligence level" drops over time. It used to drop at ~10x/year. Now it's at 30x lower in 5 months. (see image)
Assuming the $100k from before, cost could get as low as $3k/month in a few months, and if trends continue, $100/month by this time next year.
Obviously there will be smarter models and we will move to them, but replicating her experience today with equivalent models will be $1.2k/year by next year.
3. Nobody said this is normal
Lauren is the 0.0001% of "agentic engineer". Measuring her work as "$ in -> software out" is reductive and wrong.
Her spend should be looked at as a combination of product/eng, research, FDE work (pstack), and marketing. Worth $350k/day? No. Worth a mil or two a month? Absolutely.
She is exploring what happens when you ignore the cost of intelligence. She's not doing this because Cursor expects every company to spend $10m/year on tokens.
She's doing it because that $10m in tokens is going to cost $10,000 in a year or two.
Field order in your JSON schema is part of the prompt.
The model decodes left to right. If "answer" comes before "reasoning", the verdict is already sampled by the time it explains itself. That's a rationalization, not a chain of thought.
Put reasoning first and answer last.
Same model, same fields. Only one version thinks before it commits.
75% cheaper per run is the headline. The number I want is cost per solved task.
Cost per solve = price per token × tokens per attempt ÷ success rate.
A small model that takes longer trajectories or needs a retry gives the discount back fast. In agent loops it compounds, because every extra step resends the whole context.
Where it wins outright: routing, triage, and being the cheap first pass that a bigger model only has to check.
@claudeai
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.
A watermark detector answers one question: did a cooperating model sign this?
People will use it to answer a different one: is this real.
No SynthID signal means unsigned, not human. Open weights models don't sign anything.
Text watermarking nudges token sampling with a keyed score. Paraphrase enough tokens and the score drifts back to chance.
Great for provenance. Dangerous as a lie detector. @GoogleDeepMind
SynthID Detector is now available to everyone. 🌐
Check whether online content was generated using @GoogleAI, or with tools from our industry partners – including @OpenAI, @NVIDIA, Kakao and coming soon, @Apple.
Try it out → https://t.co/cPW2aNnqr6
Project Suncatcher is our moonshot exploring whether we can one day host machine learning infrastructure in space 🚀
To do so, we need to know whether our AI hardware can operate in orbit. Can our TPUs handle the physical stress of spaceflight and the radiation and thermal extremes of space? These are the questions Google researchers have been exploring for the past few years.
Last week, we launched our first test satellite carrying four TPUs into orbit — putting us one step closer to getting answers.
Pre-training went from 67% of frontier lab compute to a projected 7%. Post-training and RL went from 4% to 55%.
That changes what a GPU is for.
Pre-training is one huge synchronous job. RL is mostly rollouts: thousands of sampled trajectories, then a short gradient step.
So the bottleneck moves from interconnect bandwidth to rollout throughput.
Your RL cluster is an inference farm with a trainer attached. @vincentweisser
Unfortunately, we are being blocked by certain oligarchs in order to maintain their monopolistic chokehold on the Indian people. You can guess who they are …
This is a crime against the people of India!