This is it.
I started using Claude code recently and found that I started thinking around how can I understand the code more high level than “let me go through each of the files in detail”.
It takes a few iterations to wrap around with this approach. Once familiar, the thinking then evolves into “learn what’s needed and what’s missing and go from there aka system level”.
Sometimes it’s a major timesaver, other times it’s not that good to “let go” of control.
But overall, at the end of the day, adaptation is the best learning one can have with these tools.
This is what we've been coking for the last 9 months: make MoEs training goes ~2x faster and ~2x less memory! Highlights:
- MoE typically takes the most time and memory in modern models. Turns out one can mathematically rewrite the MoE backward pass to reduce the activation mem you need to store in the fwd by ~2x, resulting in the same gradients with no extra matmul recomputation. I really like this result, as it combines both algorithmic and systems insights.
- Analyzing bottlenecks in MoE layer leads to a natural optimization stragegy: reduce mem reads/writes as much as possible! Gathering the input for fwd and output grad for bwd can sometimes take as much time as the grouped GEMMs. We fuse gather with grouped GEMM + overlap mem access and compute to make the whole layer goes ~2x faster.
- Computing top-k for expert routing can take surprisingly long, ~15-20% of the whole MoE layer! Standard top-k impl uses radix top-k algo, great for large k but suboptimal for small k. We rewrote top-k using bitonic top-k algo, and it's sometimes 20-30x faster than pytorch's top-k!
All the main kernels are written in Cute-DSL so they should be easy to extend (and install :D). Hopper kernels are out, Blackwell kernels are just about ready. MoE models used to be 2x less hardware-efficient to train, hopefully Sonic-MOE will change that.
Going from big to small makes an inference engineer healthy and wise. Looking forward to exploring this more.
Huge appreciation to the SGLang team. @sgl_project
How long have you been "planning to understand" how modern LLM inference works?
We just gave you a readable version of SGLang you can finish over the weekend.
Introducing mini-SGLang ⚡
We distilled SGLang from 300K into 5,000 lines. Kept the core design, cut the complexity. Without sacrificing performance — nearly identical to SGLang online.
It is built for engineers, researchers, and students who want to see how inference really works and learn better from code than papers.
⭐ Star us on GitHub: https://t.co/Nk5NCXXdTz
🧵 (1/3)
We've announced cuTile, a tile programming model for CUDA!
It's an array-based paradigm where the compiler automates mem movement, pipelining & tensor core utilization, making GPU programming easier & more portable.
I'm proud of my stellar team for all their hard work on this!
Part that is interesting: CUDA isn’t the moat.
The moat is Nvidia’s fully integrated ecosystem.
AMD has had CUDA-porting tools for ~2 years (HIP) now. Google’s building their own (how good?)
Porting isn’t the hard part.
The real hard aspect is how efficiently an accelerator converts electrons into tokens.
Nvidia is still the best at taking a megawatt of power and turning it into the highest token throughput. And that advantage comes from the entire stack—hardware → NVLink → compilers → kernels → runtimes → system software. Holistic integration of components. They mostly work well once followed their support matrix.
Even if every CUDA kernel were portable tomorrow, the gap would still be:
Who produces the most tokens per watt, per rack, per dollar?
Right now, that’s still Nvidia. However, It’s promising to see silicon diversity increasing and expanding.
@hongzhou__lin Great work! best math 1.5B numbers I’ve seen! 🔥 Curious how it behaves with an extra SFT (LoRA) stage on top of DAPO, and would love to see majority@64 results too. RL usually drags majority@64 performance down, but 72.5 on AIME24 is excellent 🚀
Thanks for reminding the probability chain rule! The equivalence between token and sequence objectives is nice - it means we can think about the problem at whatever level is most convenient for analysis while knowing the optimization is fundamentally the same. I wonder how this further links to GSPO 🤔
XLA, the compiler behind JAX/TPU, once had a subtle precision pitfall: xla_allow_excess_precision=true.
It meant ops written in bf16 could silently be upcast to fp32. Intended for “safety,” but it created surprising real-world effects.
Two parallel cases showed how:
• In OSS (JAX GitHub #23543): tensors cast to bf16 were gathered in fp32, then cast back. Harmless at first glance, but this wasted bandwidth and defeated bf16’s efficiency.
• In production (Anthropic): approximate top-k + bf16/fp32 drift caused cutoff vs. selection to disagree. Edge cases meant the true top-1 token could disappear—at temperature=0, the model literally lost its most probable output. This meant your model token outputs were bad!
Why this happens:
Take logits [2.000, 1.999, …].
– In bf16, they may round together.
– In fp32, the tiny gap is preserved.
If one stage runs in bf16 and another in fp32, each “sees” a different winner. In distributed top-k or all-gather, those small mismatches get amplified across chips.
Anthropic’s final fix:
– Rolled back risky code
– Replaced approximate with exact top-k (now fast enough on modern TPUs)
– Standardized critical probability ops to fp32 (locking correctness over drift)
– Coordinated with XLA team on a compiler patch
The bigger lessons:
– Compiler defaults aren’t neutral; they actively shape outcomes.
– Mixed precision needs discipline and testing across shapes, batches, and configs.
– Distributed collectives magnify small numeric differences into large correctness failures.
– For correctness-critical paths (like sampling), accept performance cost: lock precision and use exact algorithms.
Reflection:
This is not about blame—it’s about how infrastructure complexity shapes ML workloads entirely. A single flag deep in XLA altered how models chose tokens. And because Anthropic + JAX surfaced these issues openly, the ecosystem learned together.
Some takeaways:
Prioritize exact algorithms and explicit precision control for quality-critical tasks (e.g., model sampling).
Don’t rely solely on compiler defaults or fast approximate ops when correctness matters.
Frequent validation helps quite well.
Small bugs can surface weirdly in prod! Nothing new, but helps during debug :)
https://t.co/baicCb6f3T
15K+ tokens/sec. 512-way concurrency. On just 1 GPU.
A tiny 1.5B model… doing big things. 🧠
⚡ Small model. Huge throughput. One.
Why not imagine what’s possible when a 1.5B model delivers massive intelligence per parameter?
🧠 Wouldn’t it be cool if…
…you could serve a reasoning-capable model on a single H200 GPU
at 512-way concurrency, achieving:
8.39 inferences/sec
15.6K tokens/sec
~29 ms inter-token latency
~1.3 s time-to-first-token
~60 s end-to-end on 2K-in / 1.8K-out sequences
…and even with 8K-token prompts, it still delivers:
5.45 inferences/sec
~9.8K tokens/sec ⚡
— while scoring on real benchmarks (Pass@1 avg-of-64):
🧮 AIME24 — 59.42
🧠 AIME25 — 49.68
🌍 GPQA — 42.00
📊 HMMT25 — 27.86
📘 HLE — 5.22
📚 MMLU-PRO — 55.49
➗ MATH500 — 93.80
💡 LCB — 34.50
And all that in just 1.5B parameters — tiny, fast, and cost-effective to serve.
⚡ Good news — now there is one.
Introducing (1.5B)
An open-source model designed for high-throughput, low-cost inference on real workloads.
Tuned for max throughput — 10,382 tokens/sec per billion parameters —
and it delivers.
💡 A model that’s smart enough to matter and small enough to run anywhere.
TL;DR for skimmers:
A 1.5B reasoning model that can excel in benchmarks just arrived from WRITER, it pushes 15K+ tokens/sec at 512-way concurrency — while posting strong Pass@1 scores on tough benchmarks.
Now available open source for commercial and non-commerical use :
🔗 https://t.co/28xPGyH6AI
Excited to have been a part of this work!
here is sora, our video generation model:
https://t.co/CDr4DdCrh1
today we are starting red-teaming and offering access to a limited number of creators.
@_tim_brooks@billpeeb@model_mechanic are really incredible; amazing work by them and the team.
remarkable moment.
Announcing FlashAttention-2! We released FlashAttention a year ago, making attn 2-4 faster and is now widely used in most LLM libraries. Recently I’ve been working on the next version: 2x faster than v1, 5-9x vs standard attn, reaching 225 TFLOPs/s training speed on A100. 1/