Super excited to shall our recent work! We did not cherry-pick for this "cherry picking with RL" paper!😂 Huge thanks to all my collaborators, and especially @xkelym🧡!
Let’s do 🍒 Cherry Picking with Reinforcement Learning https://t.co/4nGIJiCSBg
- 🥢 Dynamic fine manipulation with chopsticks
- 🤖 Only 30 minutes of real world interactions
- ⛔️ Too lazy for parameter tuning = off-the-shelf RL algo + default params + 3 seeds in real world
We’re releasing OmniReset, a framework for training robot policies using large-scale RL and diverse resets for contact-rich, dexterous manipulation.
OmniReset pushes the frontier of robustness and dexterity, without any reward engineering or demonstrations.
Try the policies yourself in our interactive simulator! https://t.co/3hW3nYx2vD
(1/N 🧵)
Pretrained diffusion/flow policies are powerful — but brittle at deployment.
We introduce RFS, a data-efficient RL framework that:
• steers latent noise for global adaptation
• applies residual actions for precise local correction
Works in sim and real-world dexterous manipulation 🖐️🤖
👉📄 Paper + videos: https://t.co/HumWkk7MdL
Imitation learning is great, but needs us to have (near) optimal data. We throw away most other data (failures, evaluation data, suboptimal data, undirected play data), even though this data can be really useful and way cheaper! In our new work - RISE, we show a simple way to *use all of this non-optimal data to robustify imitation learning* with minimal requirements beyond BC.
Key idea: use non-expert data to learn how to *recover* back to expert data with a minimal frills offline RL that works under sparse data coverage. Allows usage of *all* available data, not just expert data - never throw your data away!
Paper: https://t.co/gmP2V92DBL
Website: https://t.co/yi7fwPz4wi
A 🧵(1/10)
How can we create a single navigation policy that works for different robots in diverse environments AND can reach navigation goals with high precision?
Happy to share our new paper, "VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation"!
📜 Paper: https://t.co/XmyuBnrM1D
🌐 Website: https://t.co/Jt80tySWzQ
Punchline: World models == VQA (about the future)!
Planning with world models can be powerful for robotics/control. But most world models are video generators trained to predict everything, including irrelevant pixels and distractions. We ask - what if a world model only predicted the semantic information necessary for decision-making?
Introducing Semantic World Models (SWM). Given an observation and an action sequence, SWMs cast modeling as answering textual questions about the future outcome resulting from the actions. Recasting world modeling as a VQA problem lets us directly leverage the pretrained knowledge and machinery of VLMs for generalizable modeling. We had a lot of fun thinking about how this work helps connect these two seemingly very different fields of study - VLMs and world models! 🧵(1/6)
Paper: https://t.co/KIrRG2JO1a
Fun demo: https://t.co/leogQBvcO0
How can we help *any* image-input policy generalize better to visual and semantic variations?
👉 Meet PEEK 🤖 — a framework that uses VLMs to decide *where* to look and *what* to do, so downstream policies — from ACT, 3D-DA, or even π₀ — generalize more effectively!
We will have few presentations, posters and live demo tomorrow. Come hang out and chat with me at anytime~
Oral 1 & Spotlight 1 & Poster 1 & Demo (day1)
I will be on an island in the Puget Sound this weekend, so sadly I will be missing #CoRL2025tv! But luckily the amazing students who did all the work anyways, will be 😄 Here's what the WEIRD lab at the University of Washington has going on at CoRL this time
We'll be presenting 3 papers at the main conference:
1. Steering Your Diffusion Policy with Latent Space Reinforcement Learning https://t.co/fha3oTg1dx (Oral, Nominated for Best Paper)
2. ATK: Automatic Task-driven Keypoint Selection for Robust Policy Learning https://t.co/5GOqrqzwGQ
3. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies https://t.co/IA33nZZt7d (Oral)
I will be giving a talk at the RemembeRL workshop https://t.co/BdXuuoa765
Plus we have several more at the workshops! Find more details on each paper below 🧵 (1/9)
🚨Tired of binary pass/fail metrics that miss the bigger picture?
🤖Introducing #RoboEval — an open benchmark that shows *how* robot manipulation policies behave and *why* they fail, not just *if* they succeed.
🧵1/n
🔗 https://t.co/oyu7j1dzwL
📄 https://t.co/J52RLSH9Hi
So you’ve trained your favorite diffusion/flow based policy, but it’s just not good enough 0-shot. Worry not, in our new work DSRL - we show how to *steer* pre-trained diffusion policies with off-policy RL, improving behavior efficiently enough for direct training in the real world!
DSRL retains nice exploration from the base policy, but allows for quick improvement beyond this base policy with RL. The method is frustratingly simple, and super easy to throw on top of your favorite pretrained policy (VLA/diffusion policy, etc).
https://t.co/EyMsnZMCuy
Let’s think about how it works, 🧵 (1/10)
How can we continuously improve large pretrained behavior policies when 0-shot performance is not good enough? Directly finetuning the base policy via RL tends to be sample-inefficient. Can we squeeze more juice from the base policy to enable automatic and efficient performance improvement? Check our new paper, DSRL, for more details.👇
Diffusion policies have demonstrated impressive performance in robot control, yet are difficult to improve online when 0-shot performance isn’t enough. To address this challenge, we introduce DSRL: Diffusion Steering via Reinforcement Learning. (1/n)
https://t.co/ER94GwXvja
How should a robot perceive the world? What kind of visual representation leads to robust visuomotor policy learning for robotics?
Policies trained on raw images are often fragile—easily broken by lighting, clutter, or object variations—making it challenging to deploy policies learned via imitation learning in high variability test conditions. This same fragility is also reflected in the difficulty in transferring visuomotor policies from simulation to reality for robotic manipulation.
Introducing ATK https://t.co/W8fcMO6Xui: an automatic task-driven method for selecting flexible keypoint-based visual representations that enables robust, generalizable robotic manipulation with minimal human effort.(1/8)👇
Huge thanks to @shubham_kernel, Zhengyu Zhang, @xkelym, @siddhss5, @abhishekunique7, and everyone who helped along the way! 🙏
📄 Paper:https://t.co/U6uwDgbcWh
🌐 Website: https://t.co/qwyrS8GPaF
Read the paper to see what makes it tick—lots of subtle design choices and insights packed in.
💡 Takeaways:
1️⃣ Use robust yet flexible visual representations from RGB images. Keypoints are one such representation
2️⃣ Automatically select keypoint representations that are minimal and task-specific using task driven gradients from supervised learning
3️⃣ Get policy robustness and generalization for free, with minimal requirements beyond standard imitation learning assumptions. (8/8)