🤖 Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from?
📼 Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention.
🧠 VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library?
🧵 Video2Skill: From Streaming Experience to Reusable Embodied Skills
Why it matters:
📈 It can label much more data for training agents.
🔁 It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new?
📄 Paper: https://t.co/hrKsEI1WZE
🌐 Project: https://t.co/1QsXmGeTrb
🤗 Data: https://t.co/vlKWhQflC5
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone.
Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
Surprisingly, embodied data is now scaling fast: teleop, UMI, human video.
Yet embodied models aren't scaling with it.
The bottleneck now is accurate labeling 🧑💻
Our new work turns long human & robot videos into fine-grained actions, grounded to objects, in an open yet consistent vocabulary. Raw footage → reusable skills.
Long context doesn’t solve long-horizon reasoning if the model doesn’t know what to look for.
The challenge is not simply remembering more history — it’s recovering the right history at the right time. That distinction seems increasingly important as embodied agents take on longer and longer tasks.
🚨 VLAs struggle on long, context-dependent tasks, so dense progress, a score at every step, matters for them.
🤔 Can progress reward models label dense progress for these context-dependent tasks?
🧭 Our work shows why they fail: not because they are blind, but because they get lost in the context. Given the right context, the same progress reward models are reliable again.
🧵 So we propose ProgressCompass, a training-free agentic system: off-the-shelf progress models + VLMs, guiding progress estimation in long tasks.
📄 ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
🌐 https://t.co/VLHUZEZ9lZ
Thanks for sharing! Yeah VLAs now have their “memory” components, we argue that embodied Progress Reward Models also need that for dense progress reward annotation✅
Equipping VLAs with a 'memory compass'! This work introduces a training-free agentic system that combines off-the-shelf progress models with VLMs, effectively solving the core challenge of progress estimation in long-horizon, context-dependent tasks.
💡 Takeaway: a PRM doesn't need to interpret the whole history itself. It needs to be told where in the task it is. For long tasks, the right context matters more than more history.
📝 Limits: we study robot manipulation so far, and the loop relies on the VLMs reading the scene correctly.
👥 Joint work with Keliang Wu (co-first), Chengxuan Qian (@chengxuanqian), Xiyuan Yang, @zhangce1203, Ariel Tian, Anbang Liu, Haoran Lu, and Han Liu (Northwestern, UCSB, UIUC, CMU).
🌐 https://t.co/VLHUZEZ9lZ
🎬 2-min walkthrough below.
📊 Results. Wrapped in the loop, the same frozen RoboMeter goes from MAE 25.4 → 9.3 (−63%) and rank agreement 0.53 → 0.93 (+76%). It is the best method on every context form, ahead of every frozen PRM and of R²VLM.
🛡️ It stays robust when the video stops early, does more than asked, or follows an unrelated task.
⚡ Parallelizing the loop, so that no component sits idle, cuts wall-clock time by 65.6%.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
🤖 Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from?
📼 Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention.
🧠 VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library?
🧵 Video2Skill: From Streaming Experience to Reusable Embodied Skills
Why it matters:
📈 It can label much more data for training agents.
🔁 It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new?
📄 Paper: https://t.co/hrKsEI1WZE
🌐 Project: https://t.co/1QsXmGeTrb
🤗 Data: https://t.co/vlKWhQflC5
💡 Takeaway: training teaches models to use the skills they already have. The hard part is knowing when none of them fits and a new skill is needed. That is how an agent keeps learning new skills from what it sees. We think this is the key problem to solve next.
🙌 Our data is open. We'd love to see what you build!
🤗 https://t.co/vlKWhQflC5
🌐 https://t.co/1QsXmGeTrb
📝 Limits: our skills describe what happens in the video. They are not code a robot can run. We cover two domains so far.
👥 Joint work with @zhangce1203 (co-first), Xiyuan Yang, Chenwei Xu, Haoran Lu, @Williamiumli, @xie_yaqi, Katia Sycara, and Han Liu (Northwestern, CMU, UIUC, UCSD).
🎬 2-min walkthrough below.
🧱 Finding 3: trained models often reuse skills, but rarely add new ones.
🥔 Example: someone shapes a potato-chicken mix into a ball. No skill fits yet. A trained joint model finds the moment, but calls it grasp(object, source): its skill for picking up an onion.
📊 This is common. Trained joint models reuse a skill for 97–99% of repeated actions, but add a new skill for only 19–40% of first-time actions. So the library stops growing: 40 kitchen videos show 46 kinds of actions, but the library ends with about 21 skills.