First project of my PhD is out!
I'm super excited about this direction! There are plenty of unsolved problems around inverse-graphics tasks to be explored.
Cudos to my co-leads @Pkulits@ruihong04.
Super thankful for all the advising from @jiajunwu_cs
Do coding agents understand the dynamics of the world well enough to reconstruct it from video?
We introduce 4DCodeBench to evaluate this ability through 4D inverse graphics.
https://t.co/QlkzTPZpbi
<🧵>
Learning bimanual dexterous manipulation from a single human video? Agents can be a solution!
Introducing DexAgent, an agentic Human2Sim2Robot framework for dexterous manipulation with a self-evolving tool library.
We evaluated DexAgent on 11 real-world, long-horizon dexterous manipulation tasks spanning rigid, articulated, and deformable objects.
🤖 63.6% policy rollout success rate, compared with 18.2% for the strongest baseline
⚡ 2.1 hours average inference time with the tool library, compared with 3.3 hours for the baseline
🧰 A self-evolving library with 103 skills + 188 verifiers, designed to keep growing as DexAgent encounters new tasks and objects
Learn more and contribute to the growing tool library:
Project Website: https://t.co/BNRZihjElC
Arxiv: https://t.co/hao8B4btBX
So exciting to finally meet all you guys in person! So many legends turn out to be quite lovely. Feeling so supporting each other, like a family. Looking forward to seeing you again.
SGNinfy (#CVPR2023) automatically constructs 3D avatars from sign language (SL) videos. 3D avatars can increase accessibility and aiding SL learning. The challenge, however, is to accurately capture the body, face, and hands without a complex and expensive capture system.