Super excited to share Dyna-2, the first robot foundation model trained on over 1 million hours of human data. At this scale, we saw the emergence of a cross-embodiment transfer scaling law: training on increasing amount of human data not only improves model prediction on held-out human data but also on robot data the model has never seen before ๐คฏ
We validate that this transfer scaling law does translate to on-robot performance across 3 different robot platforms, and uncovers that both training objective and data matter greatly for this emergence. Crucially, our results position video as a new scaling axis for physical AI.
Beyond these scaling-law oriented results, we also dissect Dyna-2 greatly, going in-depth on its many capabilities, including language following via world modeling, enhanced robustness and precision, zero-shot customer site deployment, and some cool video generation results!
Read our technical blog post for more details! I do think this is a very important result that shows a very different path for robot foundation models from the ones we are marching on. Really excited for the road ahead. We have more releases coming, stay tuned!
Just watched The Odyssey and I have a bone to pick with Nolan. You put a sex scene in Oppenheimer but not the movie about having a string of affairs with nymphs for a decade? What the hell?
Autonomy isn't binary. Neither is robot learning.
In our latest post, CTO & co-founder @Ben_Burchfiel explains why deployment, remote assistance, and high-quality training data are deeply connected.
https://t.co/wKMBeUerjr
โAnnotationsโ are a loaded term. Every company has different needs, different techniques, different strategies. And itโs difficult to have conversations that arent grounded in data. Thatโs why weโre sharing these annotations.
Play around with the demo here: https://t.co/jJprFZT8MO
Today, we are releasing over 415 hrs of annotations on top of the expansive ABC 130K data set from @xdof. We took 5 tasks and ran them through the Shotwell engine, and here are the results: https://t.co/5zVQ0BAVH5
Automated temporal segmentation, dense descriptions of the segments, and scene descriptions of the data, with a bar for quality only humans can hit. Play around here: https://t.co/WYMhkiBkdE
We are always looking to collaborate: if youโre a research or academic group looking for annotations for your research, let us know or upload your data today: https://t.co/8QEfO9qV0k
Today, we're introducing @shotwellst.
We're solving the hardest problem in robotics: understanding your data
Thereโs no point collecting robotics data if you don't understand whats happening in it, but your options are limited. If you need dense, accurate labels with precise subtask boundaries, VLMs don't cut it, an in-house annotation army is a huge operational sink, and today's vendors charge an arm and a leg for low quality work
Shotwell combines the best of both worlds by training our own annotation models and letting humans deal with the ugly edge cases.
Its the ONLY way to consistently guarantee 100% quality in a scalable way. Humans are inconsistent, and models are inaccurate.
@Ultraroboticsco Weโre already processing thousands of hours a week for our customers today. If you have data you want us to look at, drop it here and we'll respond within the next 24 hours https://t.co/3mx4l7uA7o
@Ultraroboticsco Weโre already processing thousands of hours a week for our customers today. If you have data you want us to look at, drop it here and we'll respond within the next 24 hours https://t.co/3mx4l7uA7o
People think weโre named after the street in the mission but itโs actually the other way around. The street was named after us!
Excited to announce @ShotwellSt today :)