Politics isn’t just about policies – it’s about people. When we listen, we shape a future that reflects all of us, not just the loudest voices. Let’s debate with empathy, not just rhetoric. #Politics#ListenUp
🎉 SGLang v0.5.15 is out!
We spent this cycle tuning GLM-5.2 NVFP4 for production serving, now hitting 500+ tok/s/user on 8x B300 and 450 on 4x GB300 (bs=1).
We will put commands to run this at the thread below, and full technical details and instructions on a blog very soon 🫡
And we have some newly supported models: Hunyuan 3 (Hy3), Hierarchical Reasoning Model (HRM-Text), NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, and Qwen3.6.
Here are highlights for this release:
- Breakable CUDA Graph is now the default capture path
- Native web search built in, powered by @ExaAILabs
- Decode context parallelism for MLA models, including DeepSeek V3
- FlashInfer all-to-all for routed MoE
- DeepSeek-V4: FlashMLA sparse prefill now on by default (>10% throughput on long context), plus a non-paged indexer for long-context prefill (>5% e2e)
We welcomed 43 new contributors, and thanks again for our amazing partners and model makers: @NVIDIAAI@AMD@intel@Zai_org@TencentHunyuan@Alibaba_Qwen@deepseek_ai@Sapient_Int
Now. MAX LOAD! MAX OUTPUT! 🚀