Today, we release the first series of our findings.
1 - We cut pretraining costs by 62% at frontier scale (8x Chinchilla), and we further cut all AI inference costs by 30% (Faster decode).
2 - Introducing our model ‘Feather’, which decisively overtakes Qwen 3 as the leading frontier 1.7B math model, using 180x fewer total training tokens, even matching Qwen 4B at potent benchmarks in math.
3 - With Anvil II, our state of the art LLM Optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
The website : https://t.co/fG1TZecjIE
Our model Feather 1.7B is on air. Compare performance benchmarks through our API : https://t.co/uAJYC5VqZ6
Context : Feather is the leading math and reasoning model on earth at the 1.7B scale, beating Qwen3 1.7B using 180x less total training tokens. We further match Qwen 4B, a larger model, on potent math benchmarks using 414x fewer total training compute.
https://t.co/uhuH2yLNak
Excited to have some of my pretraining results published!
1. We’re announcing ANVIL III, the new pretraining optimizer that builds on ANVIL I and ANVIL II (the optimizers I created for the nanoGPT speedrun). ANVIL III takes a new approach, and is much more suited to large-scale models and overtraining, where the advantage grows as we scale up. It features several changes that @Everule2 and I made after the publication of ANVIL II in the nanoGPT world record. ANVIL III, in our internal testing, shaved even more time off the nanoGPT speedrun, but we are not willing to publish the IP at this moment, as it is an optimizer that can make a massive impact on the frontier of pretraining. A SNR distribution curve of ANVIL III compared to muon is attached. ANVIL III also has remarkably consistent optimal parameters across model scales and overtraining regimes, making it exceptionally well-suited to frontier pretraining.
2. Over the past 4 days, I used ANVIL III and architectural changes I had come up with back in July, to train a LLM primarily on web data. These combined to create Feather-1.7B, a 1.7-billion parameter base (only pretrained) model, trained entirely on 200B tokens of public datasets (decontaminated thoroughly across all benchmark test sets). At 180x less compute, it outperforms Qwen3-1.7B, the frontier base model at its size, on math, while matching it on reasoning. It is competitive with Qwen3-4B on math, despite being trained with 414x less training compute.
We are unfortunately compute-constrained, so we did zero parameter tuning at 1.7B scale and did only one full training run of this model. Despite likely being far from its optimized frontier, we still outperform Qwen3-1.7B and compete with Qwen3-4B on math benchmarks, without losing any language/reasoning capabilities.
3. Last but not least, the nanoGPT speedrun world record I created and submitted, taking the record from 73.889 to 39.914 seconds. This record marks a 1.851x speedup, the largest ever and larger than the last 45 world records combined.
Really excited to continue the work, and looking forward to some of the other insane results we will be publishing soon @hyperstition_cc!
Link to pretraining blog: https://t.co/PyfvnSSwVC
Excited to have some of my pretraining results published!
1. We’re announcing ANVIL III, the new pretraining optimizer that builds on ANVIL I and ANVIL II (the optimizers I created for the nanoGPT speedrun). ANVIL III takes a new approach, and is much more suited to large-scale models and overtraining, where the advantage grows as we scale up. It features several changes that @Everule2 and I made after the publication of ANVIL II in the nanoGPT world record. ANVIL III, in our internal testing, shaved even more time off the nanoGPT speedrun, but we are not willing to publish the IP at this moment, as it is an optimizer that can make a massive impact on the frontier of pretraining. A SNR distribution curve of ANVIL III compared to muon is attached. ANVIL III also has remarkably consistent optimal parameters across model scales and overtraining regimes, making it exceptionally well-suited to frontier pretraining.
2. Over the past 4 days, I used ANVIL III and architectural changes I had come up with back in July, to train a LLM primarily on web data. These combined to create Feather-1.7B, a 1.7-billion parameter base (only pretrained) model, trained entirely on 200B tokens of public datasets (decontaminated thoroughly across all benchmark test sets). At 180x less compute, it outperforms Qwen3-1.7B, the frontier base model at its size, on math, while matching it on reasoning. It is competitive with Qwen3-4B on math, despite being trained with 414x less training compute.
We are unfortunately compute-constrained, so we did zero parameter tuning at 1.7B scale and did only one full training run of this model. Despite likely being far from its optimized frontier, we still outperform Qwen3-1.7B and compete with Qwen3-4B on math benchmarks, without losing any language/reasoning capabilities.
3. Last but not least, the nanoGPT speedrun world record I created and submitted, taking the record from 73.889 to 39.914 seconds. This record marks a 1.851x speedup, the largest ever and larger than the last 45 world records combined.
Really excited to continue the work, and looking forward to some of the other insane results we will be publishing soon @hyperstition_cc!
Link to pretraining blog: https://t.co/PyfvnSSwVC
For Context :
1 - OpenAI’s stargate is spending 500bn on GPUs, with each Pretraining run, individually costing about $4bn. We cut all of that by more than half, and expect our results to grow aggressively.
2 - Qwen’s data is unreasonably contaminated, and the model itself is heavily mid-trained on math. By all means its an unfair comparison, and we still demolish them. Not only is Feather the worlds best 1.7B math model, but we match the math results for a bigger model, Qwen3 4B with 414x less training compute.
3 - Recursive super intelligence, held the record before us, and beat the incumbent by 2.2seconds. They’ve raised 650million from Google Ventures. Teams celebrate 0.8second leaps. You can only contextualise what a 34 seconds jump really means.
The next drop will be public later this month.
Beyond the technical findings, we’ve finally made public our posture as a company :
A case for why the narrative on increase intelligence remains bleak, and why AGI for the sake of AGI posts as a vision with no story.
So much more to come within the month.
We’re aggressively challenging what it takes to push the frontier.
Today, we release the first series of our findings.
1 - We cut pretraining costs by 62% at frontier scale (8x Chinchilla), and we further cut all AI inference costs by 30% (Faster decode).
2 - Introducing our model ‘Feather’, which decisively overtakes Qwen 3 as the leading frontier 1.7B math model, using 180x fewer total training tokens, even matching Qwen 4B at potent benchmarks in math.
3 - With Anvil II, our state of the art LLM Optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
The website : https://t.co/fG1TZecjIE