Learning bimanual dexterous manipulation from a single human video? Agents can be a solution!
Introducing DexAgent, an agentic Human2Sim2Robot framework for dexterous manipulation with a self-evolving tool library.
We evaluated DexAgent on 11 real-world, long-horizon dexterous manipulation tasks spanning rigid, articulated, and deformable objects.
🤖 63.6% policy rollout success rate, compared with 18.2% for the strongest baseline
⚡ 2.1 hours average inference time with the tool library, compared with 3.3 hours for the baseline
🧰 A self-evolving library with 103 skills + 188 verifiers, designed to keep growing as DexAgent encounters new tasks and objects
Learn more and contribute to the growing tool library:
Project Website: https://t.co/BNRZihjElC
Arxiv: https://t.co/hao8B4btBX
Check out RoboRender! our new work on video generation for Sim2Real robot data. We trained our own video model to turn simulation into real-world robot videos, so policies trained entirely in simulation transfer zero-shot to the real world.
Website: https://t.co/vwZKL4CD4E
Joint work with @RavenHuang4, @wensi_ai, @jwang633, Zijian Du, Yang Liu, Jiaolong Yang, @drfeifei, and @jiajunwu_cs. Thanks to all coauthors!
Sim2real visual gap has always been a challenge in robotic learning. Can video generation models be a solution?
Introducing RoboRender — a generative video framework that transforms simulated robot trajectories into realistic training data, enabling zero-shot transfer to the real world.
Across 8 tasks and 3 robot embodiments, RoboRender achieves 71% real-world success — 7.1× over raw simulation.
Real2sim has made impressive progress with coding agents. But what about sim2real, especially the visual gap? Meet RoboRender: turning simulated trajectories into photorealistic videos for policy training, achieving 71% average success in zero-shot real-world robot deployment.
Real2sim has made impressive progress with coding agents. But what about sim2real, especially the visual gap? Meet RoboRender: turning simulated trajectories into photorealistic videos for policy training, achieving 71% average success in zero-shot real-world robot deployment.
Mobile manipulation policies can break from base pose errors of just a few cm, which are common after navigation. Can we get pose generalization without additional demos? Introducing MobileVISTA: collect demos at one base pose, get a policy that works from many 🧵👇
Website: https://t.co/VKqttW2yMD
Paper: https://t.co/7SVHJeJCWP
Do coding agents understand the dynamics of the world well enough to reconstruct it from video?
We introduce 4DCodeBench to evaluate this ability through 4D inverse graphics.
https://t.co/QlkzTPZpbi
<🧵>
Teaching robots dexterous skills from human video is difficult when objects are articulated, deformable, or physically diverse. DexAgent addresses this challenge with an adaptive Human2Sim2Robot framework that turns a single egocentric human video and task prompt into physically grounded robot trajectories for policy training.
The system reconstructs the task in simulation, selects or develops task-specific tools, verifies each stage, and uses feedback to correct errors before they propagate. It can also vary object and robot states to generate diverse training data, while retaining newly developed skills and verifiers for future tasks. Across 11 real-world tasks, policies trained on DexAgent-generated data achieved a 3.5× higher success rate than competing baselines.
Title: DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library
Project URL: https://t.co/57moVHelJ1
Paper: https://t.co/TvpFwAyjVM
Hi, that's a great question! As we discussed in our paper (results analysis section) grip force is not directly inferred from the human video. DexAgent instead reconstructs the object and task in simulation based on its inference (texture, friction coefficient, weight, etc.), finds a grasp that satisfies physical constraints (e.g., contact / force closure), and verifies it through designed stress tests. So the force is goal-driven that is enough to accomplish the task rather than the human’s ground-truth force.
This is also why force-sensitive interactions remain a limitation of the current system, as this cannot be directly observed from a single video. We discussed this further in the paper, and we think this can be an interesting future direction to go!
Learning bimanual dexterous manipulation from a single human video? Agents can be a solution!
Introducing DexAgent, an agentic Human2Sim2Robot framework for dexterous manipulation with a self-evolving tool library.
We evaluated DexAgent on 11 real-world, long-horizon dexterous manipulation tasks spanning rigid, articulated, and deformable objects.
🤖 63.6% policy rollout success rate, compared with 18.2% for the strongest baseline
⚡ 2.1 hours average inference time with the tool library, compared with 3.3 hours for the baseline
🧰 A self-evolving library with 103 skills + 188 verifiers, designed to keep growing as DexAgent encounters new tasks and objects
Learn more and contribute to the growing tool library:
Project Website: https://t.co/BNRZihjElC
Arxiv: https://t.co/hao8B4btBX
Astra already shows exciting potential for real2sim2real for solving robotics tasks. Our DexAgent harness takes it further: no task-specific robot demos, 3.9× the success rate of zero-shot Astra across 11 real-world tasks (63.6% vs. 16.4%)! Check it out!
This work was done at the wonderful @StanfordSVL with wonderful people! Also, biggest thanks to the supervision of @RavenHuang4, @jiajunwu_cs, @drfeifei, and @YunzhuLiYZ, who provided me with tremendous and constructive advice!