We swept every benchmark in MLPerf Training 6.0 as the only platform to submit on all models and frameworks.
✅ DeepSeek-V3 (671B MoE) trained in 2.02 mins on NVIDIA GB300 NVL72 systems on @CoreWeave
✅ Llama 3.1 405B trained in 7.07 min on NVIDIA GB200 NVL72 systems on @Azure
✅ DeepSeek-V3 throughput improved 1.3x in just 3 months through software alone
Extreme co-design, scale-up and scale-out networking, software optimizations, and high resiliency enable the fastest time-to-train for frontier model builders.
Full technical breakdown: https://t.co/qNF2Px3UoT
It’s time to go beyond language models.
Introducing Odyssey-2 Max, our most powerful world model yet. It materially advances the SOTA in physical accuracy.
This is a big step toward models that simulate and interact with the world in real time.
A new intelligence entirely!
@Mikesfriend@BenjaminDEKR As a non-expert in robotics I would say it’s two distinct challenges, requiring very different types of movements. But maybe what you learn from running can be applied to folding laundry?
@BenjaminDEKR I’ve learned with my @maticrobots that it’s ok if the robot completes a task a lot slower than a human would. I’m at the office all day, I just want the tasks completed before I’m home.
2026 might go down as the best movie year ever:
- The Odyssey (directed by the GOAT Nolan)
- Dune: Part Three
- The Mandalorian & Grogu
If only The Batman Part II wasn’t delayed to 2027…