We’ve uploaded a great SONIC checkpoint we’ve been using internally lately!
This release focuses on supporting finer-grained manipulation, with much better wrist precision for fine-grained tasks. It works especially well with PICO’s 5-sensor mode.
Next: Thor support, less overheating, and more GR00T + SONIC demos!
What if robots, humans, and video models could all describe action in the same language? Introducing Hydra-0🐍: a generalist world model that represents robot actions as action flow. One interface. Many embodiments, tasks, and environments.
🧵1/7
Congratulations to @yukez and the @NVIDIA team. SONIC published in Science Robotics and featured on the front page of Science.
Not the robotics front page. The Science front page. Humanoid whole-body control just became one of the most important stories in all of science.
Happy and proud our motion data could play its part.
SONIC is officially published in Science Robotics today and it made the Science front page.
We show the promise of scaling motion tracking toward natural, robust whole-body control for humanoid robots.
Huge thanks to the team, and more exciting work is on the way.
Paper: https://t.co/Pa9wy4TVJw
Code: https://t.co/ZipTq2TcIt
New in @ScienceMagazine 🎉
Really excited to share that SONIC is now published in Science Robotics!
SONIC formulates super-scaling motion tracking as the foundational task for more natural and robust whole-body control. Our learned tracker and latent space further support foundation models like VLA/WAM training.
Huge thanks to the team and collaborators who made this happen.
https://t.co/rHb7z08m0V
New checkpoints coming in today!
#ScienceResearch #ScienceRoboticsResearch
Can a world model 🌏 learn to predict how 198 different deformable objects move -- not just one rope or cloth?
That question motivated Deform360: 1,980 real-world interactions captured from 41 synchronized views with bimanual touch sensors, which becomes crucial when vision is occluded.
Accepted at #ECCV2026. 🧵
1/5
How can robots learn dexterous manipulation from human demonstrations at scale?
Excited to share CHORD: Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration.
CHORD learns from human demos by focusing not only on where contact happens, but how that contact moves the object through force and torque guidance.
This unified contact-wrench representation carries human manipulation skills across diverse behaviors, long-horizon tasks, whole-body embodiments, and real-world hardware.
We evaluated CHORD on large-scale, long-horizon, contact-rich tasks paired with human demonstrations, spanning rigid, articulated, and multi-object manipulation.
At scale:
* 82.12% average success across 1,831 tasks
* 90.77% whole-body manipulation success
* 4,739 sim-ready dexterous manipulation benchmark
* Transfer to real dexterous hands
Project page: https://t.co/cAnHaJUy6P
Tech report: https://t.co/oWMzszwrgw
Code will be released soon as part of Video to Data repo https://t.co/spdPL4Gxt6, our end-to-end pipeline for converting human demonstration videos into simulation-ready assets and physics-grounded robot training data.
Huge thanks to amazing contributors: @zhu_xinghao , Zixi Liu, Shalin Jain, Chenran Li, Milad Noori, Huihua Zhao, John Welsh, @michaelv03, Wei Liu, @TingwuWang , Xingye (Dennis) Da, @zhengyiluo, Vishal Kulkarni, @sNaema, @yukez, @DrJimFan, @bowenwen_me, @danfei_xu, @SohaPouya, @Dr_YanChang.
#Robotics #PhysicalAI #DexterousManipulation #RobotLearning #NVIDIA
Humanoid robots just got more natural. 🦾
Join our Robotics Office Hours on Wed, May 20 at 11 AM PT to explore SONIC, a new NVIDIA Research framework for training humanoid robots using large-scale human motion data and reinforcement learning in Isaac Lab.
Bring your questions and join us live with our experts.
Add to calendar: https://t.co/KvcQbaDxiq
StreamdiffusionV2がMLSys26でBest Research Paper Awardを獲得しました!!
StreamDiffusionから進化させてさらに素晴らしい研究を行い、論文にしていただいたチームの皆に感謝です!!
引用先のGitHubリンクから誰でも試せますので是非遊んでみてください。
What is missing to bring real-time motion research into AAA games and real-world robotics?
We present MotionBricks, a step toward bridging this gap with two key components:
- a single generative latent motion backbone covering 350,000+ motion skills, running at 15,000 FPS with 2 ms latency and substantially improved quality and reliability.
- a unified smart primitive interface for locomotion, object / scene interaction, with fine-grained control over generated behaviors.
Webpage: https://t.co/aJE5skUuWD
Code: https://t.co/r56D3TJ8CW
Paper: https://t.co/CtOHXnHZMv (ACM TOG / SIGGRAPH 2026)
Ever wonder how we got our humanoids to walk naturally using a motion tracker? Here is the secret:
We have a world-class, game-engine-ready, real-time motion generator running under the hood!
Graphics 🤝 Robotics
Check out the full thread here 👇
Training humanoid robots?
You need motion data. Real, high-fidelity, human motion data. And until now - there was no open dataset purpose-built for humanoid robotics.
For 5 years, we've been building the largest enterprise-grade human motion and behavior datasets for embodied AI. Our data powered breakthrough SONIC research.
Today, at GTC, with @NVIDIARobotics, we're opening a piece of it to the world.
BONES-SEED:
→ 142,200 motion capture animations
→ Up to 6 natural language descriptions per motion
→ Temporal segmentation of every action
→ Curated for humanoid robotics
→ In NVIDIA SOMA and Unitree G1 (MuJoCo) formats
From text to action. Now yours.
Go build → https://t.co/00PzoIBMWe
#NVIDIAGTC
288 hours of high-quality, text-annotated human motion data are now available! 140k motion sequences!
Do you know that a large part of SONIC's training data is now open-sourced?
Check out the dataset here 👇🏻 from our friends at Bones Studio!
Full human + G1 retargeted motion!
Stie🌐:https://t.co/ui84wZXUxC
Data💿:https://t.co/Kevt6PxQ25
SONIC training code coming VERY VERY soon!
Are you still disappointed that SONIC only supports G1?
Not anymore!
We’re excited to share that 288 hours of high-quality human motions as well as G1 motions with text labels.
Special thanks to Bones Studio.
Website 🌐: https://t.co/ICCKkS7lO2
Imagine a single policy adapting aggressively across multiple embodiments and different domains—varying friction, mass, limb lengths. Can this be done online and zero-shot, without privileged environment parameters or retraining for each new domain?
We take a step toward this goal with DADP, a diffusion-based policy for domain adaptation. DADP learns domain representations in a self-supervised manner from interaction context and integrates them into the diffusion generation process by biasing the prior distribution and re-formulating the diffusion target.
Paper: https://t.co/cSsPA1Jmfc
Website: DADP: https://t.co/T5AmbLRUGp
Code (w/ Dataset & Checkpoints): https://t.co/RgssL0QJrR
More details below.
What can half of GPT-1 do? We trained a 42M transformer called SONIC to control the body of a humanoid robot. It takes a remarkable amount of subconscious processing for us humans to squat, turn, crawl, sprint. SONIC captures this "System 1" - the fast, reactive whole-body intelligence - in a single model that translates any motion command into stable, natural motor signals. And it's all open-source!!
The key insight: motion tracking is the one, true scalable task for whole body control. Instead of hand-engineering rewards for every new skill, we use dense, frame-by-frame supervision from human mocap data. The data itself encodes the reward function: "configure your limbs in any human-like position while maintaining balance".
We scaled humanoid motion RL to an unprecedented scale: 100M+ mocap frames and 500,000+ parallel robots across 128 GPUs. NVIDIA Isaac Lab allows us to accelerate physics at 10,000x faster tick, giving robots many years of virtual experience in only hours of wall clock time. After 3 days of training, the neural net transfers zero-shot to the real G1 robot with no finetuning. 100% success rate across 50 diverse real-world motion sequences.
One SONIC policy supports all of the following:
- VR whole-body teleoperation
- Human video. Just point a webcam to live stream motions.
- Text prompts. "Walk sideways", "dance like a monkey", "kick your left foot", etc.
- Music audio. The robot dances to the beat, adapting to tempo and rhythm.
- VLA foundation models. We plugged in GR00T N1.5 and achieved 95% success on mobile tasks.
We open-source the code and model checkpoints!! Deep dive in thread:
We have seen rapid progress in humanoid control — specialist robots can reliably generate agile, acrobatic, but preset motions. Our singular focus this year: putting generalist humanoids to do real work.
To progress toward this goal, we developed SONIC (https://t.co/zOZVraFuDV), a Behavior Foundation Model for real-time, whole-body motion generation that supports teleoperation and VLA inference for loco-manipulation.
Today, we’re open-sourcing SONIC on GitHub. We are excited to see what the community builds upon SONIC and to collectively push humanoid intelligence toward real-world deployment at scale.
🌐 Paper: https://t.co/DGBP7LAvuT
📃 Code: https://t.co/WAZ1P13072