Here's something most people haven't seen yet.
Before Kolibri 1.0 there was Kolibri Origin. Internal only. Same active compute per token.
On German, it went from 46 to 71. In three months.
That's not luck. It's rigor, discipline, a lot of automation, and a team that really works together. We've built a high-performing machine, and I'm proud of every person in it.
Thank you for all the feedback since launch. We read everything, the love and the criticism, and we embrace all of it. It makes us better.
Kolibri 1.1 is coming. Faster than you think.
How do you ship a model in under two months? You build the machine that builds the model.
Automate everything you can. Measure constantly. Change your mind fast. Stay disciplined.
The team wrote it all down here.
@Ulfhedinn_@quantinine fair point. we're doing text first and doing it right. multimodal is on the list.
what's the office use case for you? scanned docs, charts, forms?
ok this one made my morning.
@quantinine squeezed Kolibri to 4-bit, 42% lighter, ran it through 10,692 agentic episodes, and it still holds up.
GGUF after 11 hours. then MLX. I can't keep up anymore with what you're all building.
so tell me. what are you running it on?
Kolibri-1, now 42% lighter ⚡
Releasing Kolibri-1-NVFP4A16-FP8: Aleph Alpha's Kolibri-1 in 4-bit, on par with the original across 10,692 agentic episodes and 24 benchmarks.
https://t.co/g3Rn49uMQl
@IlhanScheer@tfburns@letiepi@alessio__serra@resiliem@tradeqvest
Shout out to @PitNeitemeier and @alessio__serra for explaining why we built Kolibri this way, breaking down the trade-offs in compute and memory, and sharing the tools that helped guide those decisions.
Thanks for opening up the thinking behind the model.
Every architectural choice has a cost. Choose where to pay it.
@alessio__serra and @PitNeitemeier derive from first principles how design choices allocate parameters, FLOPs and KV cache, compare open models, and share some of the tooling we used to design Kolibri.
@lu_zero_@Aleph__Alpha Official release is FP8. For fewer bits, the community already has GGUF, MLX and NVFP4 builds on Hugging Face.
And yes, we heard you: people want Kolibri on their laptop. Baby Kolibri is on the list 🐣
Kolibri, day 4.
#1 trending text generation model on Hugging Face.
30 community builds.
First GGUF went up 11 hours after release.
Runs on Apple Silicon, llama.cpp, Blackwell, a single 4090.
The most downloaded version isn't ours anymore.
Built by @Aleph__Alpha. Now built on by everyone.
That's how open weights should work.
This is the evaluation I hoped someone would run. 29k graded episodes, German and English, nobody from our side involved.
The part I care about most: Kolibri reads the policy before it acts, and it doesn't get worse in German. That's what regulated customers actually ask us about.
And thanks for the weak spot. English text showing up in German sessions isn't good enough. Passing it straight to the team.
Bigger birds are a fair ask. But we picked the Kolibri on purpose. The smallest bird there is, and one of the most capable. It hovers, flies backwards, and some cross the Gulf of Mexico nonstop. Small engine, absurd range. That's the design: 3.46B active, 78B total. Stay tuned for the rest of the flock :)
My favorite Kolibri experiment so far. Not for the aim, which needs work, but for the method. No fine-tuning, no generated text to parse. You read action probabilities straight out of vLLM and the game moves. 13 choices, one batch, 39ms median. That's a pattern for real-time control far beyond games. Breakout to Doom in about a day. What's next? :)
Here's something most people haven't seen yet.
Before Kolibri 1.0 there was Kolibri Origin. Internal only. Same active compute per token.
On German, it went from 46 to 71. In three months.
That's not luck. It's rigor, discipline, a lot of automation, and a team that really works together. We've built a high-performing machine, and I'm proud of every person in it.
Thank you for all the feedback since launch. We read everything, the love and the criticism, and we embrace all of it. It makes us better.
Kolibri 1.1 is coming. Faster than you think.
How do you ship a model in under two months? You build the machine that builds the model.
Automate everything you can. Measure constantly. Change your mind fast. Stay disciplined.
The team wrote it all down here.
@konarkmodi@Aleph__Alpha 100M tokens in one weekend. Truly incredible, Konark. Thank you for making Kolibri available for everyone to try, for free. You helped make this weekend special for our team 🙏
Focus. Discipline. Delivery.
Today: Kolibri. Built end to end by our team in Germany.
We also published the full tech report. Methods, results and limitations. 189 pages.
To the team: thank you for seeing it through.
Made in Germany. Built for the world.
Small bird, fast wings, Kolibri is here.
78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.
Now the weights are yours. Run it on your own hardware, under Apache 2.0.