Today we launched Kolibri. On German National Day.
A new LLM https://t.co/wfAmMTo7IF from Aleph Alpha available under Apache 2.0.
Its been an intense few months across pre and post training to make this real. Look forward to getting adoption and feedback!
@mrzhbrtweet We developed this with our amazing team in Germany. Our tech report has all the details! We went through several model iterations including an origin concept model - learning and iterating at speed was challenging but also very rewarding!
Love the reactions to our release today! Here are some of my favorite details covering architecture, load balancing and hyperparameter transfer from developing and training Kolibri, our 78B-total, 3.5B-active MoE model:
@naorweissmann We designed the model to work for core enterprise use cases like agentic rag. Particularly in settings where efficiency is needed. You can read more in the contextual performance section of our blog https://t.co/AXEFJgjhKu
Here are two neat speed tricks we found while SFT'ing Kolibri 🧵
1: We use standard Next-k-fit online bin packing with document masking maximises useful tokens per sequence. Whats easy to miss: less padding means a larger effective batchsize, so the learning rate can go up.
Our model, Kolibri, is out! 🐦
Incredibly proud of how much we've scaled in just a few months
RL is as much an infrastructure challenge as a modeling one. We've scaled our RL environments and their infra to support 20K+ concurrent sandboxes per training run and 100K+ sandboxes overall. More details in the report
A massive engineering effort behind the scenes. Proud of what we've built!
germany just dropped a sovereign open weight model
kolibri by @Aleph__Alpha runs 3.5b of its 78b parameters per word and its math is kinda ridiculous: 96.9% on aime beats every mixture-of-experts model they tested, even 3x bigger ones. only a dense model doing 8x the work wins
anyone can run it on their own servers, it thinks in german, and in their evals it tops every open model its size in english and german
models read text in chunks called tokens, and i ran kolibri's chunker (its tokenizer) on the german constitution: it needed 15% fewer tokens than gpt-5's for the same text. "bundesverfassungsgericht" is 6 tokens for gpt-5 and 2 for kolibri. fewer tokens means cheaper, faster german, and more of it fits in what the model can read at once
how it works, simply:
1. every layer has 384 tiny specialists, and a router sends each word to 6 of them. so it thinks like a 3.5b model, but it needs the memory of a 78b one: about 78 gb, which means 2 big nvidia gpus (h100s) or 1 h200
2. most layers only look at the last 512 tokens, and every 5th layer looks at everything. that's how it can read 1 million tokens (a few thick books) without it costing a fortune
3. it reasons in german. their team found that a little german reasoning data is worse than none: the model's german thoughts go in circles and never finish. so they made about 800k german reasoning examples and gave it a lot
4. it's trained to say "i don't know". they play a game with it where parts of the documents are hidden, sometimes to help it and sometimes to hide the evidence, and it has to tell which. when it didn't know an answer it admitted it 44% of the time. qwen3.5 did 11%
where it's weaker: answering from memory, using tools over a long back and forth, and coding agents, where qwen models are ahead. and to run it you need aleph alpha's add-on for vllm, a popular open source server for running models
if you have german documents and need to keep them on your own hardware, this is a big deal. huge congrats to everyone at aleph alpha, my good friend @MichaelLHofmann included!!
i wrote up how it works, the benchmarks, how to run it and when to use it: https://t.co/hcMPFukWcv
Small bird, fast wings, Kolibri is here.
78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.
Now the weights are yours. Run it on your own hardware, under Apache 2.0.
Hot off the press: We are becoming the first transatlantic sovereign AI solution together with our partner @cohere. More talent, more compute, and more innovation power to offer trustworthy AI at the security level that governments and enterprises need – across the globe.
Today, we announce a landmark agreement with @cohere. By uniting our European research depth with global AI scale, we are building a transatlantic AI powerhouse to give enterprises control over their AI.
Learn more about our shared vision here: https://t.co/6GLIosEION
Join us for the #DEsummershow featuring live online events and exhibition showcasing the graduating students and projects from our world-class courses. Register now at https://t.co/nI18qwUrbB
Join us in our new building on 22nd March for an immersive celebration of the Imperial Design Engineering community and culture of innovation. https://t.co/2mE5a3X7qi