Four years ago, in the early days of my PhD, I was obsessed with Stable Diffusion, and with the work of @robrombach, @pess_r, @andi_blatt and their team in image generation. You could say I had one north star: take these generative techniques and bring them to robotics.
Today, it feels unreal to announce @mimicrobotics' official collaboration with @bfl_ai and to reveal FLUX-mimic, our next-generation Video-Action Model for general purpose dexterity.
The bet behind it: control reduces to visual prediction. A model that can predict how a scene will unfold has already learned the physics of manipulation. So we built our Video-Action Model architecture on top of FLUX 3, the strongest video backbone available today.
We're already testing and deploying it with manufacturing leaders like @AudiOfficial, on complex manipulation long considered impossible for conventional automation. FLUX-mimic is the culmination of years of work: the promise of true multimodality extended beyond the visual, into the domain of action. A new bet on general-purpose manipulation.
And we're just getting started.
"Uncertainty Quantification for Flow-Based Vision-Language-Action Models" received the Best Paper Award at the #RSS2026 Workshop on Diffusion for Robot Learning!
Paper: https://t.co/02hPimPacP
Thanks Kaizhe Hu @hkz222, @brenthyi, @TakaraTruong and team for organizing!
𝗗𝗼𝗻'𝘁 𝗳𝗶𝗻𝗲-𝘁𝘂𝗻𝗲 𝗿𝗼𝗯𝗼𝘁 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹𝘀. 𝗦𝘁𝗲𝗲𝗿 𝘁𝗵𝗲𝗺 𝘄𝗶𝘁𝗵 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘀𝘁𝗲𝗮𝗱, 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗯𝗮𝘀𝗲 𝗽𝗼𝗹𝗶𝗰𝘆
Modern VLAs and world-action models can perform impressive manipulation skills, but adapting them reliably to new robots and tasks remains challenging.
A natural solution is DAgger-style online imitation learning: deploy the robot, collect human corrections, and update the policy. Yet foundation models are fragile in the low-data regime, fine-tuning on a handful of interventions can improve one behavior while degrading others. Online post-training or reinforcement learning can require costly data collection and exploration, making real-world learning expensive and potentially unsafe.
In our new paper, 𝗙𝗹𝗼𝘄𝗗𝗔𝗴𝗴𝗲𝗿, we take a different approach:
𝗜𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹, 𝘄𝗲 𝗹𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝘁𝗲𝗲𝗿 𝗶𝘁 𝗳𝗿𝗼𝗺 𝗵𝘂𝗺𝗮𝗻 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝗼𝗻𝘀.
The key idea is 𝗮𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝘃𝗲𝗿𝘀𝗶𝗼𝗻: we map human corrective actions back into the latent noise space of the frozen generative policy. These latent targets train a lightweight controller that adapts the robot while preserving the original model's capabilities.
Across simulation and real robots, FlowDAgger:
📈 Learns from only 5–20 human intervention episodes
🏆 Outperforms supervised fine-tuning and latent-space reinforcement learning
🤖 Works across VLAs, diffusion policies, and world-action models
✔️ Provides reliable improvements without modifying the pretrained policy
We believe this offers a practical path toward making robot foundation models improve during deployment, learning from the way humans naturally teach: through corrections.
📄 Paper: https://t.co/nzzDCJCj8W
🌐 Project: https://t.co/zIXKWUjbGO
💻 Code: https://t.co/86Pc9lZ1B4
This project was led by my amazing colleague Michael Murray with help from Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Harshavardhan Reddy Gajarla, Galen Mullins and @Andrey__Kolobov at @MSFTResearch and @mayacakmak at @UW
8🧵
Overall, our work demonstrate that uncertainty quantification improves both failure awareness and adaptation of flow-based VLAs.
If you're attending #RSS2026, I'll present the paper at three workshops and would love to chat!
w/ @mar_baga@arkrause@angelaschoellig & others
How can Vision-Language-Action Models (VLAs) tell when they don't know?
Our new paper, "Uncertainty Quantification for Flow-Based Vision-Language-Action Models", tackles this challenge.
Paper: https://t.co/ncgtKrnjz0
Website & code: https://t.co/02hPimPI2n
Details: 🧵
CLARE is a parameter-efficient continual learning method for VLAs that injects lightweight modular adapters into selected modules and autonomously expands the model only where necessary.
Recently accepted to RA-L.
Paper: https://t.co/RslZSBTHC2
Website: https://t.co/mxB4X8ysf9
World-Action Models (WAMs) have become the second dominant recipe for robot foundation models, next to classical VLAs.
So where do they come from, and how do they compare vs VLAs?
I wrote an small overview of the WAM landscape, with some personal takes:
https://t.co/6S4gH9tWTt
To all technical students and engineers in Europe.
The US export control on LLMs is the first taste of what will become the new norm. Many people are calling for "radical measures" and that we need the equivalent of a Manhattan Project to create change in Europe. But this is not how change will happen. Change will never come top-down from a government. The state and EU can fund, but they cannot found it. That part is on us.
You are the only ones who can change this.
You are among the few people on this continent who actually know how to build foundational technology - LLMs, robotic AI, actuators from scratch, chip infrastructure, rocket engines, organoids. ETH, EPFL, TUM, École Polytechnique, KTH, Imperial and dozens more produce absurd talent every single year. And almost all of it talks itself out of building.
We finish our degrees surrounded by such an incredible average that we're sure someone is always better at [your idea] - so who are we to start? I've seen so many friends at ETH think they need to "get more experience first" and take a job at Nvidia or Google and never do anything interesting again.
Technology-driven companies aren't founded by the most qualified person. They're willed into existence by people who see what others do not and refuse to stop. The person who's "better than you" almost never does it. And as for experience, nothing will teach you how to build the thing like, well, just trying to build the thing.
Our education is a chance most of the world will never have. There are people in Europe who have to worry about getting a job. We get to worry about finding our dream job. We're able to make bets that not many people can make or afford. It's nothing anybody expects you to do, but if you want a life filled with purpose, this is a unique kind of responsibility you can choose to step up to.
So if you actually want to do something ambitious, how about changing a continent?
If you really want change, you cannot wait for others. You are one of the few people who can create it. It starts with you.
@DJiafei WAMs, which everybody is currently so hyped about, also do vision+language -> action. So, VLAs aren't dying, their backbone is just changing.
How to build VLAs according to Yuke Zhu @yukez at #ICRA2026.
Happy to see Continual Learning as the crucial final stage for deployment.
That's why we built CLARE for parmeter-efficient continual learning for VLAs without forgetting: https://t.co/mxB4X8ysf9
#ICRA 2026 in Vienna is a blast! Here's our robot #autonomously participating in the robot parade!
Check out the our #SICNav-Diffusion crowd navigation method running on the robot (published in RA-L): https://t.co/cWsc2UVpeA
@florian_shkurti@angelaschoellig
(Teleoperated) Robot parade at #ICRA2026. Wondering in which year they will not be remote-controlled anymore. The technical capabilities are there, but in this kind of environment, safety is the key challenge.
Great to see our #CRISP Cartesian controller in action at the @duatic_ag booth at #ICRA2026! If you want to deploy your VLAs/WAMs on real robots, check out CRISP!
Website: https://t.co/PVvccPE33A
Code (Python & Gym APIs): https://t.co/ZiYKP8MMKq
Paper: https://t.co/j58fMeVo5e