Full-length paper on this work now on arXiv - [SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales](https://t.co/NB1PP8N4TY) - with more details on both the algorithms + systems
- Unified study of AdamW, Muon, and SOAP on dense/MoE models. Muon and SOAP consistently outperform AdamW at the scales we tested and remain stable/token-efficient at larger global batch sizes. We also identify and fix SOAP’s large-batch instability; KL-SOAP has a slight edge over Muon.
- Optimized layer-wise distributed implementation in Megatron-LM that balances memory, hides communication, and preserves the full optimizer computation: matrices assigned round-robin across DP ranks, variable-size all-gather overlapped with the forward pass, and TP-aware Newton–Schulz that stays mathematically exact.
Code: https://t.co/2suZNW0dXs
Try it out!
with @Gradientdinner et al.
Muon is having its moment — Kimi K2, GLM 5, and now DeepSeek V4!
More broadly, it feels like the time for advanced optimizers is finally here — reiterating that they are an important component for efficient training systems at scale!
Our recent work: performant Muon/SOAP-class optimizers in NVIDIA NeMo/Megatron-Core — layer-wise distributed optimizer, TP-aware Newton-Schulz, SYRK kernels. Muon ≥ AdamW on GB300-NVL72.
https://t.co/Klq27RpLAj
Muon is having its moment — Kimi K2, GLM 5, and now DeepSeek V4!
More broadly, it feels like the time for advanced optimizers is finally here — reiterating that they are an important component for efficient training systems at scale!
Our recent work: performant Muon/SOAP-class optimizers in NVIDIA NeMo/Megatron-Core — layer-wise distributed optimizer, TP-aware Newton-Schulz, SYRK kernels. Muon ≥ AdamW on GB300-NVL72.
https://t.co/Klq27RpLAj
Very excited to share the work that we've been doing on continuing to push the boundary for AI training by efficiently scaling across multiple data-centers that are over thousands of kilometers apart, using NeMo/M-Core.
https://t.co/ihXuh0w5FI
Awesome to see the Distributed Shampoo optimizer top AlgoPerf !
“28% faster training than baseline ... 19% faster than 2nd place "
Kudos to the team's tenacity for persistently improving over many months, not only surpassing strong baselines but also making it practically viable!
@MLCommons#AlgoPerf results are in! 🏁
$50K prize competition yielded 28% faster neural net training with non-diagonal preconditioning beating Nesterov Adam. New SOTA for hyperparameter-free algorithms too! Full details in our blog. https://t.co/Ge3zZ25T6D
#AIOptimization#AI
Grand Teton - Meta’s next-gen compute platform for AI !
Embodies a lot of exciting things that we’ve co-designed over past couple of years, enabling pushing our AI workloads further and beyond
#ai#ocpsummit22#codesign https://t.co/CXaUjf0rCN
Very happy share that our paper on “Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models” has been accepted for the industry track at ISCA this year :)
It’s really great to be able to showcase this work, that…https://t.co/GGPdp7rqCv
Despite being smuggled out of his hometown as a teen, @intel’s @BharatBKaul says the “universe has been kind” to him. Inspiring. #IAmIntel https://t.co/aPoQ9RcTGs
It was great to be able present one of the exec talks at the OCP global summit last week along with whitney zhao, introducing our AI training cluster and talking about the general challenges/opportunities we are seeing with buildin…https://t.co/X064RdLpx2 https://t.co/g318iP5Et2
💻@Meta’s Director of Engineering, Omar Baldonado, spoke on stage today at the 2021 OCP Global Summit, sharing the incredible work our Meta Infrastructure teams have done over the past 10 years through the Open Compute Project. Learn more about our work: https://t.co/SOEuUTh8Kl
I’m looking forward to speaking at the AI Hardware summit this year!
Perfect venue to talk about all of the interesting work on Co-designing AI HW/SW at scale at FB :)
AI HW Summit 2021 - https://t.co/vSRTS1fOls
#AIHWSummit#codesign https://t.co/bPdNouGMTo
Our work on pushing the state-of-art for training DLRMs, enabling efficient training of models with upto 12 Trilion parameters ! https://t.co/yXR7r6M1SI
NVIDIA GTC 2021 is a week-long event, that shares breakthroughs in AI, data center, accelerated computing, healthcare, gaming technology and more.
This year, GTC 2021 is hosting over 50 different #PyTorch-related sessions. Learn more below: https://t.co/o2AGzmSK0P