What is not covered in the report is how we scaled model training to be efficient on 768 B200 GPUs.
I explain it in detail here: https://t.co/MtjwZIQ7AD
A how to guide to scale up model training without wasting your GPU hours. Even if you only have 1 or 2 GPUs and cannot scale, our optimisation philosophy can still make your GPUs go brrr.
I couldn't explain ALL the details of our method in the blogpost, since the big launch came after. You can connect the dots now, the 30B MoE architecture was in fact Kolibri Origin all along.
One of my favourite things to see here is how for DeepSeek at short seqlens we read the attention state more than once (dashed read line above solid total) due to cross-layer reuse. Past 16K, sparse attention takes over, so we read it less than once and the lines cross.
What is not covered in the report is how we scaled model training to be efficient on 768 B200 GPUs.
I explain it in detail here: https://t.co/MtjwZIQ7AD
A how to guide to scale up model training without wasting your GPU hours. Even if you only have 1 or 2 GPUs and cannot scale, our optimisation philosophy can still make your GPUs go brrr.
I couldn't explain ALL the details of our method in the blogpost, since the big launch came after. You can connect the dots now, the 30B MoE architecture was in fact Kolibri Origin all along.
@m_newhaus@Aleph__Alpha Happy you liked it 🙌 a lot more optimisations had to be pulled off to get to 35% MFU in the first place, but our approach never changed, even through model iterations!
How many of your active parameters go to attention, and how many to the experts? In our new blog we derive how modern architectures allocate parameters and FLOPs, and what that means for attention state and long-context FLOPs.
Also: cute kolibris!
here we share some of the tools we used to shape Kolibri’s architecture.
they help us understand the trade-offs before committing to choices that become expensive at scale.
thought the community might enjoy playing with them too :)
@radke149 That’s true, Olmo’s dataset was very important also for us. But I doubt we’ll open source the corpus :/ if this changes, I’m sure you’ll see it publicly announced!
Kolibri, day 4.
#1 trending text generation model on Hugging Face.
30 community builds.
First GGUF went up 11 hours after release.
Runs on Apple Silicon, llama.cpp, Blackwell, a single 4090.
The most downloaded version isn't ours anymore.
Built by @Aleph__Alpha. Now built on by everyone.
That's how open weights should work.
We gave @Aleph__Alpha ‘s Kolibri-1 13 choices and put it in Doom. What could go wrong? 🎮
Up to 72 action combinations. Four decision groups, one batch.
39ms median. 60ms p95 in our test.
Apparently enough time to make questionable choices 🫣
Come judge its aim @IlhanScheer@PitNeitemeier@alessio__serra
👇 https://t.co/HV3MWejJgg
@radke149 As for data corpus, I don’t think there is a plan to open source that. Sorry to disappoint 😬
I believe the report shows what we trained on in detail, so you could reproduce it if you have the compute, or draw inspiration and scale to whatever your compute can handle
@radke149 Custom trace analyzer was just a drop in replacement for perfetto.ui so readers could interact with the trace instead of having a screenshot. Claude was very helpful there 😌 but we use perfetto.ui to analyse traces in prod