Introducing GEN-1.
Our latest milestone in scaling robot learning.
We believe it to be the first general-purpose AI model to master simple physical tasks.
99% success rates, 3x faster speeds, adapts in real time to unexpected scenarios, w/ only 1 hour of robot data.
More🧵👇
I’m so tired of writing rebuttals to this kind of “lack of novelty” review: “This paper trivially combines A, B, and C, so the algorithmic novelty is limited.”
Technically, most (if not all) robotics papers are convex combinations of existing ideas.
I still deeply appreciate A+B+C papers—especially when they deliver:
- New capabilities: the “trivial combination” unlocks behaviors we simply couldn’t achieve before
- Sensible & organic design: A+B+C is clearly the right composition—not some arbitrary A′+B+C′
- Nontrivial interactions: careful analysis of the dynamics, coupling, or failure modes between A, B, C
- Rehabilitating old ideas: A was dismissed for years, but paired with modern B/C, it suddenly works—and teaches us why
- System-level & "interface" insight: the contribution is not any single piece, but how the pieces talk to each other
- Scaling laws or regimes: identifying when/why A+B+C works (and when it doesn’t)
- Engineering clarity: making something actually work robustly in the real world is not “trivial”
- New problem formulations: sometimes the real novelty is in the reformulation—only under this view does A+B+C make sense.
Maybe worth keeping these in mind when reviewing the next A+B+C paper : )
SONIC is now open-source!
Generalist whole-body teleoperation for EVERYONE!
Our team has long been building comprehensive pipelines for whole-body control, kinematic planner, and teleoperation, and they will all be shared.
This will be a continuous update; inference code + model already there, training code and gr00t integration coming soon!
Code: https://t.co/7u3SBxzXU9
Docs: https://t.co/HpDLkTCSMF
Site: https://t.co/D3i4KlnLLr
We believe robots need instinct, not only reasoning.
Introducing Project-Instinct — a full-stack, instinct-level whole-body control toolkit for legged & humanoid robots.
🔗 https://t.co/F7xRVdUrxP
(1/3)
• Edge-Aware Safety: Novel volumetric edge penalization prevents slipping on terrain edges.
• Extremely dynamic: Robust traversal at up to 2.5 m/s!
• Open-sourced: All codes for training, deployment are open-sourced!
👇 Check out the project page!
https://t.co/Xs4iqy6qAu
Using attention-based map encoding, we enable robots to traverse extremely sparse terrains. We are now developing a new architecture that explicitly plans contact points, rather than latent representations, enabling whole-body contact planning for humanoid robots. Stay tuned.
We just released results for our newest VLA from Physical Intelligence: π*0.6. This one is trained with RL, and it makes it quite a bit better: often doubles throughput, enables real-world tasks like folding real laundry and making espresso drinks at the office.
I'm very excited to finally announce one of the most ambitious projects we've worked on — which makes the front cover of Science Robotics today:
☀️ Learning a Thousand Tasks in a Day ⭐️
Everyday tasks — like those below — can now be learned from a single demonstration each...
How do you give a humanoid the general motion capability? Not just single motions, but all motion?
Introducing SONIC, our new work on supersizing motion tracking for natural humanoid control.
We argue that motion tracking is the scalable foundation task for humanoids. So we "supersized" it: 9k+ GPU hours and 100M+ motion frames.
But tracking alone is not enough; we show how to make a useful control system out of it:
- Universal Kinematic Planner: Enables game-like gamepad control and high-level teleoperation, just like controlling a character in a game.
- VR Full-Body Teleop: Direct, real-time whole-body control by a human wearing a VR headset.
- VR Keypoint Teleop: Control the upper body (hands/head) while our planner handles robust locomotion automatically.
- VLA Integration: We connect this motion tracker to autonomous Visual-Language-Action (VLA) models for autonomous task execution!
We use a Universal Token Space to UNIFY this command space, turning our robust tracker into a general-purpose, programmable humanoid brain.
This is the generalist "System 1" for humanoids. 🚀
Project: https://t.co/X5xl7daKAS
#Humanoids #Robotics #AI #FoundationModels #NVIDIAResearch 🧠🔥
Meet BFM-Zero: A Promptable Humanoid Behavioral Foundation Model w/ Unsupervised RL👉 https://t.co/3VdyRWgOqb
🧩ONE latent space for ALL tasks
⚡Zero-shot goal reaching, tracking, and reward optimization (any reward at test time), from ONE policy
🤖Natural recovery & transition
Excited to introduce TWIST2, our next-generation humanoid data collection system. TWIST2 is portable (use anywhere, no MoCap), scalable (100+ demos in 15 mins), and holistic (unlock major whole-body human skills).
Fully open-sourced:
https://t.co/fAlyD77DEt
What if robots could improve themselves by learning from their own failures in the real-world?
Introducing 𝗣𝗟𝗗 (𝗣𝗿𝗼𝗯𝗲, 𝗟𝗲𝗮𝗿𝗻, 𝗗𝗶𝘀𝘁𝗶𝗹𝗹) — a recipe that enables Vision-Language-Action (VLA) models to self-improve for high-precision manipulation tasks.
PLD couples real-world residual reinforcement learning with standard supervised fine-tuning — letting robots discover, recover, and distill their own data flywheel.
Quick 🧵
Simulation drives robotics progress, but how do we close the reality gap?
Introducing GaussGym: an open-source framework for learning locomotion from pixels with ultra-fast parallelized photorealistic rendering across >4,000 iPhone, GrandTour, ARKit, and Veo scenes!
Thread 🧵