Congratulations to the @Dyna team! Quite an impressive read.
The finding we'd push on is the second one.
Action-labelled data and unlabelled video improve robot performance on separate axes, and footage that fails a hand-pose bar can still be exactly right for world modelling. That has an awkward implication for how most pipelines are built.
If you have one validator tuned to your labelling requirements, everything it rejects disappears, including the material that was never meant to be labelled in the first place. Same footage, two different definitions of usable, and only one of them is being applied.
We'd argue the gate has to know which pile it's filling.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
It began as a fix for a problem inside one lab. It became the way the world talks to itself.
To the idea that started it and the people who never stopped building on it.
Here’s to three and a half decades of HTML.
@lukas_m_ziegler Wall-E walked so OctoBot could wrap itself around your Amazon package. Curious to know what the training dataset behind it looks like.
You don't need someone to say "I'm scared" to know they're scared.
A shaky voice.
Long pauses.
Faster breathing.
Unlike humyns, voice AI usually cannot pick on these signals.
A thread on how it learns.
Physical AI models don't break just because of compute or model size only, sometimes they break because cameras alone can’t teach a machine how the world actually feels, sounds, or moves.
I mean, Sight was only ever the starting point. So right now, @humynlabs is solving the real bottleneck for physical AI by building The Multi-Sensory Platform for Physical AI. They're fusing sound, sight, motion, and touch into one synchronized training signal.
Humyn Labs is also powered exclusively by @KGeN_IO's multi-million verified human network across 20+ countries in different regions. Mannn, this is the real-world infrastructure layer Physical AI actually needs.
So now, I think...
Vision-only training was just a temporary workaround. Fused multi-sensory data is the actual destination.