Can robots learn a bimanual dexterous skill from just one video?
DexImit makes it possible ๐โโ๏ธ โ automatically converting human videos into physically plausible training data for bimanual manipulation, either from text2video models or real-world recordings.
1/ We just released ฯ0.7 โ a steerable generalist robot model with emergent capabilities.
I want to share a bit of the backstory, because ฯ0.7 taught me something surprising about where robot learning is heading. A thread on bittersweet lessons ๐งต
Introducing GEN-1.
Our latest milestone in scaling robot learning.
We believe it to be the first general-purpose AI model to master simple physical tasks.
99% success rates, 3x faster speeds, adapts in real time to unexpected scenarios, w/ only 1 hour of robot data.
More๐งต๐
Can robots learn a bimanual dexterous skill from just one video?
DexImit makes it possible ๐โโ๏ธ โ automatically converting human videos into physically plausible training data for bimanual manipulation, either from text2video models or real-world recordings.
๐ Result
We evaluate DexImit along two dimensions:
input video quality and task difficulty.
DexImit can generate diverse short-horizon data using only a text2video model. For long-horizon tasks such as beverage making and apple cutting, higher-quality videos are necessary.