Do you think that robot safety is “just collision avoidance”? In the open world, safety is more than collisions and must represent failures like spilling or breaking. Our new latent safety filters detect–and prevent–any policy from violating hard-to-specify constraints! (1/10)
World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics?
Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies.
DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry.
But no single modality captures everything — how can we effectively combine them?
We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best.
Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data!
🧵 [1/8]
We argue for a new era of robot safety: one that recognizes that the safety of bits (AI safety) cannot be separated from the safety of atoms (physical safety)
We also propose a taxonomy of risks ⚠️ & full-stack research opportunities 👩🔬 to guide safe generalist robot development
What does safety mean when your robot can do “anything”? 🦾💥
A few of us have been thinking about this Q for the past year and wrote a position paper:
Rethinking Safety for Generalist Robots
https://t.co/DXzmg31UHX
w/@RohanSinhaSU, @anushridixit111, @tianran_, @Majumdar_Ani!
Robots with generalist capabilities require us to rethink safety beyond collision avoidance and alignment. 💥 Check out our new position paper, with a taxonomy of emerging risks and research directions:
https://t.co/qvVG7gElo3
Gemini Robotics 2 comes with a new open benchmark for agentic safety reasoning (ASIMOV-Agentic): refusing unsafe tasks, quantifying uncertainty, and asking for human help.
Check out the safety tech report for more:
https://t.co/9WMBE73epJ
Great effort to crowd-source robot manipulation failures! I've had to explicitly collect a lot of failure data for various projects, and it's always been a pain 😅. This allows researchers (and their robots) to learn from others' failures!
This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as ✨OopsieData✨, please join us!
Pretrained ViTs see the world in rich, dense detail. Most policies pool it to a single vector before acting, discarding most of it.
We introduce Patch Policy: a minimal architectural extension that enables transformer-based policies to consume dense tokens directly, no billion-param VLM required. It outperforms a fine-tuned 7B VLA by 18% with ~0.7% of its parameters, enabling robust, precise manipulation.
As general-purpose robots edge closer to deployment, the ultimate question remains: Should we build safety into the architecture at the source, or catch failures at runtime?
Join us to hear from the incredible lineup of Invited Speakers at our RSS workshop this Friday!🤖✈️
I’m presenting Uncertainty-Aware Policy Steering at #RSS2026 in Sydney 🇦🇺 this week!
⏰ Imitation Learning 2 session on Wednesday, July 15 at 3:30pm
Feel free to reach out if you’d like to chat about this work or anything else - super happy to connect!
I'm attending #RSS2026 for a couple papers/talks and as an RSS Pioneer! Please reach out to chat and/or come to the following:
- Monday 9:40 AM: Spotlight pres on detecting side effects at the Rethinking Safety Workshop by @ryanlindeborg (https://t.co/RNZR3P7owE)
- Tues 4-5PM: Come visit my RSS Pioneers Poster! @RSSPioneers
- Wed 3:15PM: Robometer Reward Fxn talk + poster after by @yigitkkorkmaz and @aliangdw (https://t.co/nd1jpgRM0C)
- Thurs 3:15PM: TMRL pre-training for post-training talk + poster after by @matthewh6_ (https://t.co/goYhXBw7TZ)
- Friday 9:30AM: Giving an invited talk at SemRob Workshop (https://t.co/aseCY9NRCj)
- Friday 5:30PM: Invited talk at the Diffusion Workshop (https://t.co/jpKl1zmTWz) on TMRL
Diffusion / flow-based robot policies unlock two axes of test-time scaling: sequential (denoising steps) and parallel (samples). Both improve performance but cost latency on a robot, and knowing which to scale a priori is often unclear. ELASTIC learns it via RL!
1/ Inference-time policy steering lets robots imagine possible futures and pick the best one—but what if the best-looking future still feels wrong?
In ViTaL, we argue that for contact-rich tasks, steering robots by using visual futures 👁️ 🌍 is not enough. Robots should also verify what they expect to feel 🤚🌍 as a result of their actions.
This is a normal day in our lab, but don't tell our advisor! Simulation provides a way to avoid such expensive behaviors, but the consequences of actions, like fluid damage or breakage, aren't observable in sim, making it hard to train and evaluate robot policies for safety.
🧵👇
How do we make robot policies robust to rare but high-impact failures?
Video #World#Models (WMs) are rapidly becoming a powerful tool for robotics, enabling policy evaluation and improvement by "imagining" future outcomes. But there's a catch: these imagined futures are typically nominal samples, making it easy to overlook the rare yet safety-critical events that matter most.
In our new paper, StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement, we explore a simple but powerful idea:
💡 Instead of passively sampling futures, actively steer world model imaginations toward high-impact yet still plausible scenarios.
StressDream optimizes the initial diffusion noise at inference time, allowing us to generate targeted stress-test scenarios without retraining the world model. This enables:
- More robust policy evaluation by exposing failure modes that random sampling often misses.
- Improved policy optimization by training against challenging but realistic imagined futures.
As generative world models become a foundation for #Physical #AI, the ability to systematically probe their "long tail" of plausible futures will be increasingly important for building reliable and trustworthy autonomous systems.
📌 𝖯𝗋𝗈𝗃𝖾𝖼𝗍 𝖯𝖺𝗀𝖾: https://t.co/0BSM0d0X0n
📄 𝖯𝖺𝗉𝖾𝗋: https://t.co/JQ3Ij4SZWY
Work led by Junwon Seo, with a great set of collaborators: Sushant Veer, Thomas Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Andrea Bajcsy.
@NVIDIADRIVE@NVIDIAAI
#Robotics #WorldModels #PhysicalAISafety #AISafety #AutonomousSystems #RobotLearnin
Do we still need Uncertainty Quantification for robotics? What does uncertainty mean in the modern robotics paradigm?
💡 We are organizing a workshop “Rethinking Uncertainty for Modern Robotics Paradigms” @ IROS 2026
Submit your work by Aug 23! 👇 (https://t.co/AUhtTyDmnK)
Interaction with the real world is the major bottleneck in robot learning. So what would robot RL look like if we didn’t need to limit compute per interaction? Our latest work, Off-Policy Generative Policy Optimization (OGPO, accepted to ICML26) embarks on answering this question (spoiler alert: when done correctly, it helps massively!).
🧵(1/N)
Excited to share the first paper of my PhD!
If you’ve ever tried to control a VLA via natural language, you know it rarely does what it is told. 🗣️ We introduce a multi-stage pipeline for training a Language Feedback Policy (LFP) to steer a VLA in-the-loop.
Check out Junwon’s new work on how we can steer video models to generate plausible high-impact (e.g., failure) outcomes at inference time! Super important for robust policy evaluation / policy improvement in tail scenarios.
Video world model imaginations🌎💭can miss critical but plausible outcomes of robot actions.
Introducing 𝙎𝙩𝙧𝙚𝙨𝙨𝘿𝙧𝙚𝙖𝙢: inference-time steering for video WMs, imagining plausible✅, high-impact⚠️ futures for 𝙧𝙤𝙗𝙪𝙨𝙩🛡️ policy evaluation and improvement.
(1/15)
Video world model imaginations🌎💭can miss critical but plausible outcomes of robot actions.
Introducing 𝙎𝙩𝙧𝙚𝙨𝙨𝘿𝙧𝙚𝙖𝙢: inference-time steering for video WMs, imagining plausible✅, high-impact⚠️ futures for 𝙧𝙤𝙗𝙪𝙨𝙩🛡️ policy evaluation and improvement.
(1/15)
So excited to share 𝚆𝙴𝙰𝚅𝙴𝚁 🌎, co-led with @arnavkj95 !
World models are becoming a powerful tool for robot learning — but for real robots, they need to be more than visually realistic. They also need to be consistent, efficient, and useful for decision-making. 🤖✨