Punchline: World models == VQA (about the future)!
Planning with world models can be powerful for robotics/control. But most world models are video generators trained to predict everything, including irrelevant pixels and distractions. We ask - what if a world model only predicted the semantic information necessary for decision-making?
Introducing Semantic World Models (SWM). Given an observation and an action sequence, SWMs cast modeling as answering textual questions about the future outcome resulting from the actions. Recasting world modeling as a VQA problem lets us directly leverage the pretrained knowledge and machinery of VLMs for generalizable modeling. We had a lot of fun thinking about how this work helps connect these two seemingly very different fields of study - VLMs and world models! 🧵(1/6)
Paper: https://t.co/KIrRG2JO1a
Fun demo: https://t.co/leogQBvcO0
I will be on an island in the Puget Sound this weekend, so sadly I will be missing #CoRL2025tv! But luckily the amazing students who did all the work anyways, will be 😄 Here's what the WEIRD lab at the University of Washington has going on at CoRL this time
We'll be presenting 3 papers at the main conference:
1. Steering Your Diffusion Policy with Latent Space Reinforcement Learning https://t.co/fha3oTg1dx (Oral, Nominated for Best Paper)
2. ATK: Automatic Task-driven Keypoint Selection for Robust Policy Learning https://t.co/5GOqrqzwGQ
3. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies https://t.co/IA33nZZt7d (Oral)
I will be giving a talk at the RemembeRL workshop https://t.co/BdXuuoa765
Plus we have several more at the workshops! Find more details on each paper below 🧵 (1/9)
Proud to be part of this work on ATK!
Robots don’t need to see everything—just the right things. But how do we teach them what matters amid lighting changes, distractors, and ever-changing scenes?
Here’s how we solve it with ATK👇
How should a robot perceive the world? What kind of visual representation leads to robust visuomotor policy learning for robotics?
Policies trained on raw images are often fragile—easily broken by lighting, clutter, or object variations—making it challenging to deploy policies learned via imitation learning in high variability test conditions. This same fragility is also reflected in the difficulty in transferring visuomotor policies from simulation to reality for robotic manipulation.
Introducing ATK https://t.co/W8fcMO6Xui: an automatic task-driven method for selecting flexible keypoint-based visual representations that enables robust, generalizable robotic manipulation with minimal human effort.(1/8)👇