We have just released #Alpamayo 2 Super — @nvidia’s frontier open reasoning model for #autonomous#vehicles.
NVIDIA Alpamayo 2 Super is an open 34-billion-parameter #reasoning vision-language-action (#VLA) model designed to accelerate autonomous vehicle (AV) development. It combines the 32-billion-parameter NVIDIA #Cosmos 3 Super Reasoner with a 2-billion-parameter diffusion-based Action Expert and is post-trained with #reinforcement #learning.
Two aspects make Alpamayo 2 Super particularly exciting:
1. Open and commercially deployable: Alpamayo 2 Super is available on Hugging Face under OpenMDW-1.1, the Linux Foundation’s permissive license for open AI model distribution. The OpenMDW license is now being applied across the entire Alpamayo model family, enabling developers to deploy these models commercially without requiring additional permissions.
2. A multi-task foundation model for autonomous driving: Alpamayo 2 Super produces five tightly coupled outputs:
- A trajectory describing the vehicle’s planned path.
- A chain-of-causation (CoC) trace explaining the reasoning behind each driving decision, achieving benchmark-leading reasoning performance at frontier scale.
- A meta-action (e.g., yield, lane change, stop) capturing the model’s intent.
- Reasoning auto-labels that generate CoC annotations for training and validation data.
- Visual question answering responses with 2D visual grounding, linking answers to specific regions in camera images.
These multi-task capabilities enable developers to leverage a single foundation model across more of the development process, simplifying tooling and accelerating iteration.
Resources:
🔹 Interview: https://t.co/PDpu08RoZK
🔹 Video: https://t.co/lLXESgpuAn
🔹 Blog: https://t.co/cW3lb0CFtD
🔹 Technical blog: https://t.co/DCOjE2fepI
🔹 Hugging Face blog: https://t.co/iD202NTCz4
🔹 Model weights: https://t.co/xjYX2C1kT7
🔹 Inference code: https://t.co/fO9T1iNESV
As @JensenHuang has emphasized, open models help advance safety and security. We hope Alpamayo 2 Super will contribute to this vision by enabling researchers and developers around the world to build, experiment, and innovate in autonomous driving.
We are excited to see what the community builds with Alpamayo 2 Super.
@NVIDIADRIVE@NVIDIAAI
Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles.
Beyond seeing, Alpamayo understands and reasons through the complex world - thinks before it acts.
It’s a powerful backbone for robotaxis, trucks, shuttles, delivery vans, tractors and the long tail of mobile robots—billions of autonomous machines someday.
We’re releasing it for commercial use under OpenMDW-1.1 so teams can inspect it, fine-tune it and deploy it—open models advance safety and security.
The next wave of AI is robotics—and it starts with autonomous vehicles.
Great work, Alpamayo team!
https://t.co/2PYCCXWjZh
How do we make robot policies robust to rare but high-impact failures?
Video #World#Models (WMs) are rapidly becoming a powerful tool for robotics, enabling policy evaluation and improvement by "imagining" future outcomes. But there's a catch: these imagined futures are typically nominal samples, making it easy to overlook the rare yet safety-critical events that matter most.
In our new paper, StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement, we explore a simple but powerful idea:
💡 Instead of passively sampling futures, actively steer world model imaginations toward high-impact yet still plausible scenarios.
StressDream optimizes the initial diffusion noise at inference time, allowing us to generate targeted stress-test scenarios without retraining the world model. This enables:
- More robust policy evaluation by exposing failure modes that random sampling often misses.
- Improved policy optimization by training against challenging but realistic imagined futures.
As generative world models become a foundation for #Physical #AI, the ability to systematically probe their "long tail" of plausible futures will be increasingly important for building reliable and trustworthy autonomous systems.
📌 𝖯𝗋𝗈𝗃𝖾𝖼𝗍 𝖯𝖺𝗀𝖾: https://t.co/0BSM0d0X0n
📄 𝖯𝖺𝗉𝖾𝗋: https://t.co/JQ3Ij4SZWY
Work led by Junwon Seo, with a great set of collaborators: Sushant Veer, Thomas Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Andrea Bajcsy.
@NVIDIADRIVE@NVIDIAAI
#Robotics #WorldModels #PhysicalAISafety #AISafety #AutonomousSystems #RobotLearnin
LAP now ranks 2nd on the Molmospace leaderboard @allen_ai, and is the only model that is:
1. Fully open-sourced (data, checkpoint, code)
2. Evaluated out-of-the-box, even _without_ fine-tuning on DROID!
3. From academia
https://t.co/ATZwjxzlUC
🎮 Can we learn interactive world models from letting robots “play”?
➡️ Introducing ✨PlayWorld: a framework for training high-fidelity video world models from large-scale autonomous play experience that enables:
→ Accurate dynamics prediction
→ Reliable policy evaluation
→ RL fine-tuning entirely inside the world model
🌐https://t.co/Kpd2DoveXc
Today's state-of-the-art VLAs struggle to generalize zero-shot to new robot embodiments, despite training on extensive multi-embodiment data.
We introduce Language-Action Pre-training (LAP) and LAP-3B — the first VLA to achieve substantial zero-shot transfer to unseen real-world robot embodiments, through simply aligning action representation with language.
Everything is open-sourced! Try it out on your own robot:
🌐 https://t.co/dWuDtbV0Tn
Happy to share our work 'Actions as Language' is accepted to #ICLR2026!
Key idea: use a language-based action representation to better align the robot fine-tuning data with the VLM.
This reduces distribution shift & enables generalization via LoRA alone. See you in Rio!
New #NVIDIA Paper
We introduce Motive, a motion-centric, gradient-based data attribution method that traces which training videos help or hurt video generation.
By isolating temporal dynamics from static appearance, Motive identifies which training videos shape motion in video generation.
🔗 https://t.co/TbKXjQMN3H
1/10
My group @Princeton is hiring!
We are looking for strong postdoc and PhD candidates to join our quest for intelligent robots in open-world environments. Read more below and get in touch 🤖🐅🧡
https://t.co/7o35pwPZCz
Robotic manipulation has seen tremendous progress in recent years but rigorous evaluation of robot policies remains a challenge!
We present our work: "Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators"!
🧵
Video models hallucinate a lot. However, they can't self-verbalize their confidence, unlike LLMs. Meet S-QUBED, the first method to empower video models to express their uncertainty through Bayesian Entropy Decomposition!
Interested to work on generalist robots that are safe, trustworthy, and capable? 🤖
📢 My group at @Princeton is looking for PhD students and postdocs this cycle!
PhD: please apply through the MAE department (Dec. 1).
Postdocs: please email me directly!
🏆 Congrats to the best paper winner @ CoRL '25's Eval&Deploy Workshop:
📃**Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators** by @ApurvaBadithela!
🔗Paper: https://t.co/o8IiUliDDP
Thanks to @JasonMa2020 and @DynaRobotics for sponsoring the prize!
Why don’t VLAs generalize as well as their VLM counterparts? One culprit: catastrophic forgetting during fine-tuning. 🧠
We introduce VLM2VLA: a training paradigm that preserves the VLM capabilities while teaching robotic control.
https://t.co/vq96lMW0ck
🧵
🔎Can robots search for objects like humans?
Humans explore unseen environments intelligently—using prior knowledge to actively seek information and guide search. But can robots do the same? 👀
🚀Introducing WoMAP (World Models for Active Perception): a novel framework for embodied open-vocabulary object localization that combines the reasoning power of VLMs 🧠with the physical grounding capabilities of world models 🌎.
🌐 https://t.co/e6ZUra86S1 🧵(1/N)
How can we boost the visual generalization of VLAs 𝑤𝑖𝑡ℎ𝑜𝑢𝑡 𝑎𝑛𝑦 𝑓𝑖𝑛𝑒-𝑡𝑢𝑛𝑖𝑛𝑔?
@AsherJHancock will be presenting at #ICRA2025 on Wed in the Vision-Language-Action Models session (WeDT21):
Run-Time Observation Interventions Make Vision-Language-Action Models More Visually Robust
🚨 I'm on the job market looking for tenure-track positions in robotics and autonomy🤖! I'm a PhD Candidate at @Princeton
My research focuses on interactive motion planning in the joint space of physical🌍 and information📊 states (e.g., beliefs), actively ensuring safety and improving efficiency as robots autonomously navigate uncertain environments and interact with people.
My work draws on:
♟️ Game Theory
🧠 Machine Learning/AI
🔄 Control and Optimization
Reposting is much appreciated!
More info on my website:
https://t.co/a3q4JW6eQY
My research thrusts👇
Tired of your vision-language-action (VLA) model failing catastrophically in the presence of distractions?
Check out BYOVLA: Bring Your Own VLA: a run-time intervention scheme that markedly improves performance with distractor objects and backgrounds.
https://t.co/xGTmY7XTDn