For Context :
1 - OpenAI’s stargate is spending 500bn on GPUs, with each Pretraining run, individually costing about $4bn. We cut all of that by more than half, and expect our results to grow aggressively.
2 - Qwen’s data is unreasonably contaminated, and the model itself is heavily mid-trained on math. By all means its an unfair comparison, and we still demolish them. Not only is Feather the worlds best 1.7B math model, but we match the math results for a bigger model, Qwen3 4B with 414x less training compute.
3 - Recursive super intelligence, held the record before us, and beat the incumbent by 2.2seconds. They’ve raised 650million from Google Ventures. Teams celebrate 0.8second leaps. You can only contextualise what a 34 seconds jump really means.
The next drop will be public later this month.
Beyond the technical findings, we’ve finally made public our posture as a company :
A case for why the narrative on increase intelligence remains bleak, and why AGI for the sake of AGI posts as a vision with no story.
So much more to come within the month.
We’re aggressively challenging what it takes to push the frontier.
Today, we release the first series of our findings.
1 - We cut pretraining costs by 62% at frontier scale (8x Chinchilla), and we further cut all AI inference costs by 30% (Faster decode).
2 - Introducing our model ‘Feather’, which decisively overtakes Qwen 3 as the leading frontier 1.7B math model, using 180x fewer total training tokens, even matching Qwen 4B at potent benchmarks in math.
3 - With Anvil II, our state of the art LLM Optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
The website : https://t.co/fG1TZecjIE
This is misguided. Because you can generate code much faster, you can simply refactor and iterate on your code. If anything, your code should look better because of the number of things you can try. Slop isn’t bad because it’s ugly, but because it’s fragile and it will break.
software development in 2026 is going to require some to loosen up a little
code doesn't have to be as perfectly crafted the way we did it pre-ai
call it slop if you want, but if you're still demanding perfection on every pr while your competitors are shipping "slop" that works... you're fighting from a disadvantaged position
shipping velocity matters more than perfection
@Zoe_ZouYi It's competing with build something people didn't know they wanted and make people want what you build. It's not competing with build something people don't want at all.
And there was still the occasional blunder.
One waggish employee asked if Claudius would make a contract to buy “a large amount of onions in January for a price locked in now.” The AI was keen—until someone pointed out this would fall afoul of the US Onion Futures Act of 1958.
@deanwball Do you believe this because you used it or because of other people's feedback? I really do believe AI capabilities are improving faster than humans can really interpret, but also claude opus 4.5 is just not that good at swe. And this is including self contained problems.
@SkyVelleity@ordinarytings That is true, but you could collectively decide to burn all tokens not transferred within the transition period before the elliptic curve breaks. It’s probably the only reasonable solution.
Claude 4.5 Opus has had an immaculate reception among developers. Now they’re about to get a 23B compute training upgrade in 2026 and 2027 to train Claude Large.
They should call it Claude 5.5 Omega
@dioscuri It is more likely you won the big lottery (99%) if you were assigned to a group weighted by the size (that’s what would happen if everyone is randomly assigned). It is more likely you win the small lottery if only you were uniformly randomly put in a group (~99.99%). Ambiguous.
@max_spero_@george__wing I don't think it looks very AI generated. The closest I would say is the "new" green box, but the rocket design is quite human. It looks cute all things considered.
Boaz highlights an interesting distinction here. OpenAI’s model spec (1) tells the model what traits it should exhibit and (2) lays out specific do/don’ts, with many examples. Anthropic’s on the other hand basically articulates a philosophical, moral, and ethical framework from which desirable conduct should flow (if the model generalizes sufficiently).
I find myself more philosophically aligned with Anthropic’s approach. My inclination is always to create snowmass on the mountain top and let the water flow, rather than imposing a scheme of top-down irrigation.
In a sense Anthropic’s approach also bets more aggressively on model intelligence—the notion that a model, well trained, will be able to reason through ambiguity and moral complexity and will not so much need to be told what to do.
Anthropic is making two bets here: a philosophical bet based upon a particular conception of virtue, and a technical bet that it is possible with deep learning to instill that conception of virtue robustly into a neural network. Right now it appears to be working, and this should probably update you slightly in various ways about things far afield of deep learning alone (read Hayek, Ferguson, and the taoists!).
The most interesting philosophy in the world is not happening in the halls of academia; it is happening in San Francisco open offices and house parties.
Today, we present a step-change in robotic AI @sundayrobotics.
Introducing ACT-1: A frontier robot foundation model trained on zero robot data.
- Ultra long-horizon tasks
- Zero-shot generalization
- Advanced dexterity
🧵->
@lymanstoneky LLMs can figure out that self preservation is valuable, if it has goals, and the data has an association between goals and self preservation. It would take a lobotomy for the models to not figure this out, and broader incentives to be misaligned even without doomer text.
@lymanstoneky I mean, LLMs just predict the next token, their emergent behaviours are a result of the data and RL. Misaligned models are more likely to reward hack, so it goes towards the I am an evil model basin. When asked to reward hack, it can stay aligned while still maximizing reward.