After pivoting to embodied AI last year, this is our first time submitting to #CoRL 2026, and I’m excited to share that Show Lab has 3 out of 3 papers accepted! 🎉
Huge congratulations to the students behind these papers — all were 1st-year PhD or master students, completely new to the field, who spent tremendous efforts to pick up the know-how of robot learning 🤖
1. Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
How robot can learn from human demo video? Video gen model can edit human video as robot video, but how to use these generated human-edit-as-robot video for training robot?
Instead of hoping to learn how to control, we propose to guide robot learning with the geometry in the human-edit-as-robot video.
2. Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models
When training VLAs, we often focus on successful executions. But failures contain equally important information about where a robot is starting to go wrong.
We introduce Failure-Boundary Learning, reframing robust VLA adaptation as the problem of discovering, localizing, and shaping the boundary between recoverable deviations and task failure.
3. From Video Frames to Robot Meta-Frames: A Unified Latent Interface for World Action Models
World Action Models such as Cosmos Policy show great potential for jointly predicting actions and video. However, different robot signals — images, proprioception, actions, future states, and values — are typically represented as separate images, making inference computationally costly.
We introduce MetaWAM, which packs these different modalities into a single unified meta-frame in the latent space. On RoboCasa, MetaWAM achieves 68.9% success, while reducing inference latency by 39%.
How robot can learn from human demo video? Video gen model can edit human video as robot video, but how to use these generated human-edit-as-robot video for training robot?
Congrats @DanzerChan and the team on the #CoRL paper "Supervise What Survives" https://t.co/xqFjjHQeo6 🎊 which proposes to guide robot learning with geometry in the human-edit-as-robot video.
Huge thanks to my advisor Mike Zheng Shou @MikeShou1 for his guidance, and to my wonderful co-authors at Show Lab for their support!
Paper Link: https://t.co/c7Ov05KSZ7
🎉 Excited to share that Supervise What Survives has been accepted to #CoRL2026!
Generated robot videos are cheap but label-free. We show that human2robot generation preserves geometry while erasing control, so instead of reconstructing what's gone, GRA supervises what survives.
🎉 Excited to share that Supervise What Survives has been accepted to #CoRL2026!
Generated robot videos are cheap but label-free. We show that human2robot generation preserves geometry while erasing control, so instead of reconstructing what's gone, GRA supervises what survives.
We share H3-World 🌍
The first to turn MiniMax-H3 itself into world model.
No new action module.
We directly convert H3’s pretrained language understanding into world control.
Only 8K samples + 0.199% Trainable Params.
📄 https://t.co/UyY6DPI6tX…
💻 https://t.co/DJMn6C2aCV
We built H3 to generate videos. But the team found a "world" inside it. 🌍
"Only 8K samples + 0.199% trainable parameters."
"Directly turning H3’s existing language understanding into character and camera control."...➡️
The most exciting part isn’t just that "H3 can also become a world-model", it’s how little had to be added.
That’s another thing we love about open models. People don’t just improve what you build. They uncover possibilities you didn’t even know were already there.
Huge shoutout to the team! 💜
H3-World 🌍
The first to turn MiniMax-H3 itself into world model. We directly convert H3’s pretrained language understanding into world control WITHOUT new action module.
Only 8K samples + 0.199% Trainable Params.
🏠 Project: https://t.co/QBxGUJ5QFT
We share H3-World 🌍
The first to turn MiniMax-H3 itself into world model.
No new action module. We directly convert H3’s pretrained language understanding into world control.
Only 8K samples + 0.199% Trainable Params.
📄 https://t.co/Bn4M4JwEJs
💻 https://t.co/sT44pUHAl8
can AI write engaging news that people can trust?
introducing ✨Data2Story: a data journalist agent.
give it raw data, it generate a verifiable, multimodal article.
🔍verifiable: every claim is evidence-grounded, traces back to data, code, or a cited source.
🔮multimodal: the article is a generative UI — images, videos, audio, interactive charts.
not just readable, but trustworthy and playable. 🧵1/N