Order something to a robot will be almost free just as intelligence is. For agentic intelligence of robots, the best way is using an agentic LLM to reason and give orders and intellectual plans of it's intensions and a neural net to do each micro task but no need to look human:
We got our robots to wash pans, clean windows, make peanut butter sandwiches, and more!
Fine-tuning our latest model enables all of these tasks, and this has interesting implications for robotics, Moravec's paradox, and the future of large models in embodied AI.
More below!
Introducing GEN-1.5, a one-shot learner.
It can learn new tasks in a few seconds. Show it what to do, and it generalizes.
This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
Playing with Gemini Robotics 2 on Apollo is so fun!
The incredible thing about whole body generalization in 3D space on humanoids is that all the recovery, retry and reactive behaviors look very human-like - like having a proto human intelligence in a metal body in your lab.
Unseen task, unscripted motions.
1M hours of human video sounds insane.
That’s roughly 170 years of continuous waking experience.
If you had people wear cameras specifically to collect it:
1,000,000 hours ÷ 8 hours/day = 125,000 person-days
That’s ~343 people collecting 8 hours a day, every day, for an entire year.
That’s a massive data operation.
So I’m really curious: how exactly did Dyna-2 collect its 1M+ hours of egocentric human video?
Congrats @DynaRobotics on new scaling result. I encourage everyone to read https://t.co/ntsVVtgUFj, a lot more details.
"More data" is a forever ask. How much is enough? With Dyna's plots, if enough = robot MSE match human MSE @ 1M, you need 37B hr 🤔 is that right? Ur thought?
Congratulations to our NVIDIA Research team for SONIC being published in @ScienceMagazine. 🦾
Built by NVIDIA Robotics researchers and collaborators, SONIC demonstrates scaling motion tracking for natural, robust humanoid whole-body control.
Read the paper 📄 https://t.co/19FcMbU59T
How well do egocentric human videos scale? Apparently, really well.
Dyna Robotics has introduced Dyna-2, a world-action model (WAM) pre-trained on more than 1 million hours of real egocentric human video. It learns by jointly predicting the next video frames and the actions, and that video prediction (world modeling) is what carries learning from humans to robots.
• More human video predictably improves action prediction.
• Task accuracy improves across embodiments (bimanual arms and a humanoid) with zero robot data in pre-training.
• Post-trained on just a few hours of robot data, mean task performance rose 20% to 53% as human pre-training scaled 1k to 1M hours.
WAM vs VLA, matched apple-to-apple:
the WAM hit 1.55x the VLA's success rate and won 65% of head-to-heads. Deployed zero-shot at unseen customer sites, it passed 87% vs the VLA's 46%.
The $11 billion company with no product, no revenue, and no timeline. On purpose.
Physical Intelligence just closed roughly $1.6B across two rounds at an $11.2B valuation. It sells nothing. It plans to sell nothing for a while. And both Jeff Bezos and OpenAI are on the cap table. Here's the bet.
The company:
- Founded 2024, San Francisco. About 50 to 100 people, deliberately small
- Known as Pi, ticker-style: their models are named pi-0, pi-0.5, pi-0.6
- Mission: build one general-purpose robot brain that works across any robot body
- Not building robots. Building the thing that makes everyone else's robots useful
The founders are the reason the money showed up:
- Karol Hausman, CEO, ex-Google Brain robotics
- Sergey Levine, UC Berkeley professor, one of the most-cited reinforcement learning researchers alive
- Chelsea Finn, Stanford professor, pioneer of meta-learning
- Plus Lachy Groom, ex-Stripe, and a bench of DeepMind alumni
- This is the robotics equivalent of the OpenAI 2015 founding team: the field's top academics, concentrated in one room
The valuation ladder, two years:
- March 2024: ~$70M seed at ~$400M, OpenAI participating
- November 2024: $400M Series A at $2.4B, led by Jeff Bezos with Thrive and Lux
- November 2025: $600M Series B at $5.6B, led by CapitalG
- Mid-2026: ~$11.2B, Founders Fund and Lightspeed joining
- Roughly $2.1B raised total. Zero revenue against all of it
What the product actually is:
- pi-0: a vision-language-action foundation model. Show it a task, tell it in plain language, it moves the robot
- The demos that made it famous: folding laundry, bagging groceries, making coffee, tasks robotics failed at for decades
- They open-sourced pi-0's weights on Hugging Face, the frontier-lab playbook applied to robotics
- The thesis in one line: manipulation will generalize the way language did, and whoever owns the general robot brain owns the layer above every hardware maker
The contrast that frames the whole sector:
- Figure AI: $39B, builds the humanoid body
- Skild AI: $14B, another brain play, SoftBank-backed
- Physical Intelligence: $11B, brain only, hardware-agnostic
- Pi is the Android bet. Figure is the iPhone bet. Both cannot be fully right
What makes this one different from every other AI story this year:
- The company states, on the record, that it has no fixed commercialization timeline. Research first
- Every markup is priced on demos, papers, and team, not customers
- Its valuation is about 5x capital raised. Figure trades at 22x. The market is pricing the capital intensity of teaching robots honestly
- The risk is symmetrical: no deployments means no proof, and DeepMind, Nvidia, and Tesla are converging on the same territory with infinite money
The strange bedfellow detail: OpenAI and Bezos are both investors. Amazon runs a million warehouse robots. OpenAI has its own robotics ambitions. Both are paying to see the general robot brain get built by someone else first.
The frame: language models got their foundation-model moment in 2020 and the world noticed in 2023. Physical Intelligence is the bet that robotics is at its 2020 right now, that the laundry-folding demo is the GPT-2 of the physical world. $11.2B says the smartest robotics researchers alive are two model generations from proving it. No revenue required. Yet.
https://t.co/5ZEWBeP0eJ
This is the most exciting moment of my life so far.
And it’s only the beginning.
Scalability is EVERYTHING.
Over the next few weeks, we’re going to unpack 6 breakthroughs that we believe are critical to scaling physical intelligence:
Foundation models ← this is the first big unlock
Infrastructure
Teachability
Deployment
Hardware
Physical agents
None of these pieces work alone. They compound.
Today, we’re starting with the foundational model breakthrough that changed how we think about what’s possible.
The rest of the story is coming.
Buckle up. 🔥
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
Jensen is back in the house!!
Jensen and Nvidia have been phenomenal partners. I'm excited to be working closely together as we massively scale up this year