Our model Feather 1.7B is on air. Compare performance benchmarks through our API : https://t.co/uAJYC5VqZ6
Context : Feather is the leading math and reasoning model on earth at the 1.7B scale, beating Qwen3 1.7B using 180x less total training tokens. We further match Qwen 4B, a larger model, on potent math benchmarks using 414x fewer total training compute.
https://t.co/uhuH2yLNak
For Context :
1 - OpenAI’s stargate is spending 500bn on GPUs, with each Pretraining run, individually costing about $4bn. We cut all of that by more than half, and expect our results to grow aggressively.
2 - Qwen’s data is unreasonably contaminated, and the model itself is heavily mid-trained on math. By all means its an unfair comparison, and we still demolish them. Not only is Feather the worlds best 1.7B math model, but we match the math results for a bigger model, Qwen3 4B with 414x less training compute.
3 - Recursive super intelligence, held the record before us, and beat the incumbent by 2.2seconds. They’ve raised 650million from Google Ventures. Teams celebrate 0.8second leaps. You can only contextualise what a 34 seconds jump really means.
The next drop will be public later this month.
Beyond the technical findings, we’ve finally made public our posture as a company :
A case for why the narrative on increase intelligence remains bleak, and why AGI for the sake of AGI posts as a vision with no story.
So much more to come within the month.
We’re aggressively challenging what it takes to push the frontier.
Today, we release the first series of our findings.
1 - We cut pretraining costs by 62% at frontier scale (8x Chinchilla), and we further cut all AI inference costs by 30% (Faster decode).
2 - Introducing our model ‘Feather’, which decisively overtakes Qwen 3 as the leading frontier 1.7B math model, using 180x fewer total training tokens, even matching Qwen 4B at potent benchmarks in math.
3 - With Anvil II, our state of the art LLM Optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
The website : https://t.co/fG1TZecjIE
We are now @hyperstition_cc.
Today, we release the first series of our findings.
1 - We cut all pretraining costs by 55% at Chinchilla optimal and 66% in overtrained regimes, and all AI inference costs by 30%.
2 - Our model Feather 1.7B beats Qwen3 1.7B on math, using 180x fewer total training tokens.
3 - With Anvil II, our previous state of the art optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
Learn more : https://t.co/4pSlwAbdbT
The Nano-GPT speedrun has been, in effect, the most potent contextualization of pre-training efficiency. On its first iteration in 2024, @karpathy set the baseline at 45 minutes. The ability to shrink that to under a minute, in a tidy couple years, stands as a testament to the pace of LLM advancement.
Our public implementation beats the current Nano-GPT speedrun record by 8.7 seconds, bringing the new world record to 65.87 seconds - the single largest absolute jump in over 18 months.
We have insofar redacted the capabilities that substantially move it to under the 60 second threshold for reasons proprietary to SPL.
This, though nimble, marks as a singular, among a panoply of findings SPL keenly expects to release over the coming weeks.
Kudos to @DevenPietrzak who individually theorized and implemented it within a few short nights, while simultaneously pitching and taking to the European finals, Poland’s premiere baseball squad.
Here’s the github PR : https://t.co/yZsnlfL635