Kimi K3 is live on Modal.
Moonshot has shipped the world's first open 3T-class model, and we're a day zero launch partner.
We trained a custom DFlash speculator for K3's novel architecture so you can run it faster, losslessly. The most capable open model we've worked with by far.
made a small artifact to explore update geometry across submissions to @kellerjordan0's nanoGPT optimization track
helped me build some intuition for what these optimizers do during training, maybe it will for you too!
Speculation Is All You Need.
In this blog post, we announce the co-release (w/ Z Lab) of six more state-of-the-art DFlash speculators for @Alibaba_Qwen 3.x.
Over 1k output tps for 3.5 122B-A10B on a B200.
Read the blog for why we're all-in on spec dec.
https://t.co/Bv3Zc95Xgh
Running 100k sandboxes is a hard scaling problem that's not for everyone – @modal is very excited to be one of a small set of providers that can handle this.
In fact, @modal was the only one to nail every single iteration of this test leading up to the final results.
We worked with @lmsysorg and https://t.co/Cg0JsVomui to
- integrate DFlash spec into @sgl_project
- make it faster with overlap
- train a DFlash drafter for @Alibaba_Qwen 397B-A17B
The result: up to 4.3x greater throughput over baseline and 1.5x over native MTP.
🚀 New blog: The next generation of speculative decoding: DFlash and Spec V2
DFlash + Spec V2 hit >4.3X baseline throughput for LLM inference, now the default speculative decoding engine in SGLang! Together with @modal and https://t.co/ZXetBKIRym, our jointly-released DFlash drafter for Qwen 3.5 397B-A17B beats both baseline and native MTP in every setting we benchmarked:
1️⃣ >4.3X baseline & 1.5X native MTP throughput (concurrency 1, HumanEval, 8xB200)
2️⃣ Block diffusion drafter: a full token block in one forward pass
3️⃣ KV injection: target-model features fed into every draft layer’s KV cache for higher acceptance
4️⃣ Spec V2 overlap scheduler: +33% end-to-end
Read the code, deploy a DFlash server, and start experimenting!
Last fall, we shared our deep dive on FA4 internals.
But we didn't stop at grokking the kernel.
Since then, we've been developing improvements for inference performance and upstreaming them.
This blog post explains those contributions.
https://t.co/xzDNHdq3Zw
Tried to squeeze the most important bits about the entire stack for cloud deployment of transformer inference, from application layer concerns to hardware, debugging, and o11y, into one talk. Had to operate at a very high tok/s!
https://t.co/CFlfGyCSOs