Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
Two and a half weeks ago Dyna Robotics published that their model is actually working, and now they are lifting the curtains on some of their deployments.
Dyna-2 is a world-action model pre-trained on 1 million hours of first-person human video, which equates to training on about 170 years of hands. They unlocked transfer. With the same checkpoints, scored on robot data the model never saw, zero fine-tuning, their scaling curve holds. They hit a human-to-robot scaling law. Somewhere between 10k and 100k hours, the embodiment gap starts closing on its own.
At a customer site neither has seen, Dyna-1 passes 46%. With Dyna-2 scaling successfully, it passes 87%, enabling Dyna to finally start scaling their deployments.
Din Tai Fung, one of the highest revenue-per-location chains in the US, is rolling Dyna robots across its network, with hundreds of robots by H1 2027.
We’ve seen plenty of impressive robotics demos. But now the bar must move higher: robots must create real value in the real world — and that value must scale.
For us, the real test is not whether a robot can complete a task once. It is whether customers get enough value that they want to deploy more. Robotics only matters when it solves real problems: taking repetitive, tedious, dirty, and difficult work off people’s hands while creating meaningful economic value for the businesses using it.
ROI and scalability are inseparable. One successful deployment can prove customer ROI, but not a scalable product or business. If every new customer or workflow requires rebuilding the solution, the economics will never scale. You only prove that ROI can scale when the same underlying technology keeps creating value across customers, workflows, and industries — while the effort, cost, and time for each new deployment keep coming down.
We think about the path very simply:
Demo: “It works.”
You’ve proven technical possibility.
Pilot: “It works here.”
You’ve proven it can work in a real environment and start creating value.
Scaled deployment: “We want more.”
Customers keep expanding because the ROI works, and we can keep delivering that value without rebuilding everything from scratch.
Every deployment must make the next one better. A problem solved in the field should leave something reusable behind — a better model, better tooling, better infrastructure, or a more general capability. What we learn from one customer should make the next deployment easier, faster, and more reliable.
This is also why research and deployment must stay tightly connected. Research expands what robots can do. Deployment tells us what actually matters, where things break, and what must be solved fundamentally — not patched case by case. You need both to build a product that can truly scale.
Today, we’re sharing more of what we’ve learned across a broad range of industry partners — the successes, failures, operational challenges, and hard-earned lessons behind getting robots to create real value in production.
The goal is not just to make robots work. They must create real value. That value must repeat. And it must scale.
Many Thursdays ago, I got this text:
“Understand u guys r quite tied up with existing deployments, but when can we get our next batch of robots?”
I read it three times.
Anyone who has built something from 0 to 1 knows the feeling. For months, you’re pushing a boulder uphill. You don’t know if there’s a top. Then one day, almost quietly, you feel the boulder start rolling on its own. The question had changed from “does this work?” to “when can we get more?”
We found our fit.
The path there was not obvious. For a long time, the question I got most about Dyna was: are you a model company or a deployment company?
What I rarely admitted was that I wasn’t entirely sure either.
We knew what we were doing, but it felt like we were swimming against the tide because investors kept pushing us to pick a lane. The cleaner path was to just build the brain. It was a story with no revenue pressure, software-like scalability, better multiples, and less messiness. But the real world is brittle, and robots are not LLMs. So we kept coming back to what we were actually here to build.
The answer wasn’t another breakthrough model. And it wasn’t more unscalable deployments. It was a robot people wanted.
It took us a while to realize there was a third path. Frontier research makes the product possible. Deployment teaches us what the research needs to solve next. Each makes the other better. Both are necessary to build a PRODUCT people want. After all, isn’t that why we’re building robots in the first place?
We are Dyna. A product company that turns frontier intelligence into real-world impact.
There’s still more to come.
An hour of lab evals catches a model that doesn't work. It won't catch one that fails once every two hundred trials, or degrades over a week, or runs fine on this robot and badly on the one beside it.
So the eval moved to where the work is. Every episode, every site, graded on the customer's definition of good. Over a terabyte a day, autolabelled into SOP steps, outcomes, and failure modes.
Today, a new Dyna deployment goes from setup to production ROI in as little as three days.
This is just scratching the surface.
Also - we're hiring across deployment research: https://t.co/Xh6wTZ5tuk
Let’s dive into the facts. The small details we obsessed about and actually closed the gaps.
To keep one Din Tai Fung dining room supplied, a robot has to produce 1,500 table-ready napkins in a day’s shift.
A year ago, ours managed 480 with Dyna-1. Today with Dyna-2, we clear well over 1,500 napkins, 2.7x faster and at higher quality. None of it came from one-off solutions engineering. And the performance gains didn’t come from the pre-training alone, but the extensive systems around it.
Then we realized there is more. A fully folded napkin is not a finished job. It has to land in a bin, stacked, ready for staff to carry out so they can serve their guests.
For months our napkins were folded well, but the bins were disorganized and ruined output quality drastically. The robot was placing each napkin wherever it happened to land, so staff had to square up the stacks before carrying them out.
Ten positions in one small bin look nearly identical to the robot. Nothing in the scene tells it what comes next, only the instruction does. Precise language following in Dyna-2 is what closed it, and it generalizes: a new bin size is a different instruction, not a different model.
Waymo was founded in 2009. Their demo took 18 months. Their first driverless rides opened to the public in 2020.
It's easy to show a demo. What comes after are the edge cases, recovery, uptime, integrations, and the small failures that only show up once a robot has done the job every day for months. In the field.
We spent the past year in that "after" with our customers, running their napkin operations day in and day out. Twelve months ago, Dyna-1 was the most reliable robot foundation model DEMO published at the time. Today the comparison with Dyna-2 is stark.
After so many demos, models, pilots, robots are still struggling to land real deployments with real customers. Until now.
Today we’re excited to share that our robots have successfully crossed the ROI threshold, and Din Tai Fung, one of the highest revenue per location restaurant chain in the US, is rolling out Dyna robots across its extensive restaurant network.
This brings our rollouts across hotels, logistics, data centers, and many other use cases to a fleet that reaches hundreds of robots by the first half of 2027. And we’re just getting started.
It’s been a wild year, and today we’re double clicking on the battlefield stories and sharing a few learnings about scaling robot deployments. We are just scratching the surface.
Read the full blog post: https://t.co/8QVEoQuzBY
I will be at #Actuate26 giving a talk on the model and infrastructure behind dyna-2. Excited to chat with everyone about scaling robot foundation models and deploying them in the real world! Please reach out if you'd like to chat!
Feeding the GPUs is one problem; what runs on them is another. We rebuilt our Muon optimizer to shard state inside each node instead of across the job, avoiding the slower inter-node fabric as node count grows:
• optimizer step time: ~3x faster at scale
• job resilience: preflight node health gating + auto-restart from last checkpoint, new clusters stood up in days instead of weeks
If working on these systems is interesting to you: https://t.co/Xh6wTZ5tuk
When people talk about robotics, they usually talk about models, data, or hardware. Few people talk about the infrastructure that lets you iterate on all three quickly. Today we're publishing how we trained Dyna-2 on over 1,000,000 hours of egocentric video, repeatably. At this scale, most of what worked at ten thousand hours did not hold up:
• ingestion throughput was capped at 14,000 episode-hours per week — a million hours would have taken over a year
• building a training manifest took 48 hours before a run could even start
• reading a petabyte from cloud storage during training left GPUs exposed to latency and packet loss
🧵
Our training clusters don't always sit next to our data, and reading a petabyte-scale corpus straight from cloud storage exposes GPUs to egress latency and packet loss. We built an on-cluster NVMe caching layer on top of Alluxio:
• time to read one petabyte (single thread): 57.9 days (cloud) → 5.8 days (cluster-local cache)
• steady-state throughput: ~2 GB/s per node per read, keeping GPUs at 98% utilization
Jason Ma reveals the first true scaling law in robotics: robots trained on 1 million hours of first-person human video predictably get better without ever seeing robot data
"Yesterday we announced our new flagship robot foundation model called Dyna-2. It's very significant for the entire field for many different reasons. First of all, it's the first robot foundation model trained on at least one million hours of data. That in itself is a very challenging infrastructure challenge."
"What we demonstrated is that even if the one million hours of data is entirely first-person video of humans doing manipulation tasks, we actually see scaling law for transferring to robot embodiments that the model has never seen before."
"By training on more human data, we actually see predictable performance improvement on robots. That's a big deal because the biggest challenge facing robot foundation models is that we don't have enough data."
"Unlike language models, there are just not enough robotics data out there, and it's very difficult to collect these robot data. Only by passively observing humans doing things can we scale this kind of data, and this is the first time we showed there is a lot of improvement that can transfer."
@JasonMa2020@DynaRobotics
Congratulations to the @DynaRobotics team. This is probably my favorite "scaling laws for robots" blog post so far. I hope that it *does not* remain my favorite, and that the community continues to raise the bar further.
Next phase: robot brain companies start to expose inference endpoints or remote "ask me anything" sessions to test out their model on their robot.
My favorite part of this blog post is that it provides enough detail that the result could be reproduced by an external lab (1M hours egocentric data is quite obtainable). They even evaluated on setups that any lab could buy (ABC-style bimanual YAM).
Robotics is entering a scale-up era. The scale of investment is very serious, and so warrants serious rigor when companies make claims about models that only they can verify. Otherwise, we risk vaporizing billions of VC dollars underwritten by self-reported evaluations of capabilities.
Indeed, I love working with Jason and our entire research team.
There was a time when people questioned us: Is this team really capable of doing impressive research? We don’t have professors. We don’t have industry celebrities. We don’t have the familiar names people immediately recognize.
I was unhappy hearing that, because I knew how strong our team was. But I stayed quiet. Because I also knew that if we didn’t ship, we couldn’t prove anything.
Today, I finally want to say it proudly: we have a world-class research team.
Their names might not be familiar to everyone yet, but make no mistake—they are rock stars. I’ve seen firsthand their talent, creativity, rigor, and determination, and I’m incredibly proud of what this team has accomplished.
In the end, it’s not about where you come from.
It’s about where we’re heading together.
Really big thanks to our awesome research team!
@JasonMa2020@kun_h____@_anhquanpham@chetan_@georgejygao@TopiwalaAnirudh@johnnywang_16@YifeiRobotics@tianyurobot
At the same time, we are still hiring more world-class researchers for all the directions!!!!!!!!! Please reach out if you want to raise this scaling law to another level!!!!!!!
Applied Researcher - Deployment Intelligence & Continuous Learning: https://t.co/iGxcnfTjXh
Research Engineer/ Scientist: https://t.co/6BjJBAtb38
Research Engineer/Scientist, Simulation: https://t.co/0JRG6C9g5K
Research Internship: https://t.co/Ovn2xOOGJC
This is the most exciting moment of my life so far.
And it’s only the beginning.
Scalability is EVERYTHING.
Over the next few weeks, we’re going to unpack 6 breakthroughs that we believe are critical to scaling physical intelligence:
Foundation models ← this is the first big unlock
Infrastructure
Teachability
Deployment
Hardware
Physical agents
None of these pieces work alone. They compound.
Today, we’re starting with the foundational model breakthrough that changed how we think about what’s possible.
The rest of the story is coming.
Buckle up. 🔥
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵