"Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding"
Speculative decoding for RL rollouts!
This paper speeds up post-training without changing the target policy’s sampling distribution. So a draft model proposes multiple tokens, and the policy model verifies them.
An important system piece is that the draft must stay aligned with the continually updated policy, with weight sync and optional online adaptation.
This gives faster rollouts, but same learning trajectory. With ~1.5-1.8x rollout speedup at 8B, and projected ~2.5x end-to-end speedup at 235B scale.
RL post-training is hitting a rollout bottleneck.
This new paper from #NVIDIAResearch shows how speculative decoding in NeMo-RL + @vllm_project can accelerate rollouts losslessly, with 1.8x higher throughput at 8B and projected 2.5x end-to-end speedup at 235B.
Read the full paper: https://t.co/twR4LEQNmy
Announcing NVIDIA Nemotron 3 Super!
💚120B-12A Hybrid SSM Latent MoE, designed for Blackwell
💚36 on AAIndex v4
💚up to 2.2X faster than GPT-OSS-120B in FP4
💚Open data, open recipe, open weights
Models, Tech report, etc. here:
https://t.co/CAYpP1iK3i
And yes, Ultra is coming!
What if you could ask a chatbot a question the size of an entire encyclopedia—and get an answer in real time?
Multi-million token queries with 32x more users are now possible with Helix Parallelism, an innovation by #NVIDIAResearch that drives inference at huge scale.
🔗 https://t.co/7vAr3IFfwx
Inference at scale is pushing the boundaries of what today’s systems can handle. One promising direction? #DisaggregatedInference — splitting the serving pipeline into distinct stages like prefill and decode — to unlock better performance across the throughput-interactivity spectrum.
While the concept has gained traction and sparked a wave of open source experimentation, real-world deployments remain rare. Why? Because the optimization space is vast, and orchestrating system-level coordination is incredibly complex.
📄 In our latest #NVIDIAResearch, we present a comprehensive systematic study of disaggregated inference at scale, evaluating hundreds of thousands of design points across varied model sizes, traffic patterns, and hardware configurations.
Key findings:
1️⃣ Disaggregation shines in prefill-heavy traffic patterns and with larger models.
2️⃣ Dynamic rate matching and elastic scaling are essential to achieving Pareto-optimal performance.
3️⃣ There's no one-size-fits-all architecture — workload-aware tuning is critical.
Whether you're building production inference infrastructure or exploring #LLM optimization techniques, this work offers actionable insights to balance system throughput with low-latency responsiveness.
🔬 Read the paper and dive into the data — let’s shape the future of scalable #AI serving, together.
📗https://t.co/TRqywTK0Px
👀New #NVIDIAResearch on boosting MoE model performance with disaggregated serving.
Learn how our NVIDIA Dynamo and GB200 NVL72 work together to boost the performance of #AI data centers running MOE models like DeepSeek R1 and the new Llama 4. ⚡
Technical deep dive➡️ https://t.co/bdAWZ0m97E
🚀 Want to improve your LLM responses? Read our tutorial for implementing AmbigNLG! Addressing task ambiguity in Natural Language Generation to drive more accurate, context-aligned outputs. @iso_map
https://t.co/OMMANV6nKw #AmbigNLG#NLP#tutorial#LLMs#EMNLP2024#MLEngineering
🌴Heading to #EMNLP2024! Presenting AmbigNLG with @ayaniwa1213 on Tuesday at 4pm (Riverfront Hall).
Paper: https://t.co/ktKO6ywnup
Data: https://t.co/tve3hZ18Ok
You can also stop by our @MegagonLabs sponsor booth or DM me to chat about full-time and internship opportunities :)
📢Excited to bring the DAIS workshop to ICDE'25 (w/ @SainyamGalhotra @FarihaAnna @MikeCafarella@sairamgv) The focus is on the emerging idea of compound AI systems with a specific emphasis on data discovery, interactions w/ data, architectures for #AgenticAI+#LLM, and evaluation.
Introducing 🫧 HoloBench—a new benchmark for measuring LLMs' reasoning capabilities across massive document collections!
- RAG models fetch info well but struggle with multi-doc reasoning.
- HoloBench evaluates how LLMs synthesize and aggregate info to solve complex tasks.
🧵 1/
We hope our work sets a new standard for evaluating holistic reasoning in LLMs! Proudly co-authored with @SAYg_7 (equal first author!) and Nikita Bhutani at @MegagonLabs.
💻 Code: https://t.co/Ox7vpnOJZt
🗂 Dataset: https://t.co/P0FiS87F8p
6/
⚡️Are you working with #NLP or #AI-driven products for Natural Language Generation (#NLG)? Task ambiguity is a common pain point and AmbigNLG changes that! AmbigNLG is designed to solve task ambiguity in instructions for NLG. What it is... 👇🧵
🎉 Our long paper has been accepted for the #EMNLP2024 main conference!
In "AmbigNLG," co-authored with @iso_map, we tackle task ambiguity in NLG instructions to better align LLM outputs with your expectations.
📃 Read more: https://t.co/DQ7a4sseaT