Very excited to introduce Cosmos 3: Omnimodal World Models for Physical AI! 🚀
https://t.co/xwrFwHsXY6
Cosmos 3 works across language, image, video, audio, and action . It brings together capabilities that often live in separate systems: multimodal reasoning, image/video/audio generation, action modeling, world simulation, and robot policy learning, all within a single unified omnimodal world model 🌎. This creates a more direct path from perception to simulation to control 🤖.
Cosmos 3 also ranks #1 among open models on multiple reasoning and generation benchmarks and leaderboards. This was a huge one-team collaboration across NVIDIA. It brought together research, engineering, data, simulation, infrastructure, deployment, and many other efforts. I am deeply grateful to everyone who contributed and proud to have been part of this journey.
Technical report: https://t.co/5qBo1Js5yi
Models: https://t.co/qBi2j9gvkH
Code: https://t.co/ZlIeDxW3lx
Website: https://t.co/xwrFwHsXY6
#nvidia #cosmos #physicalai #worldmodels #robotics
This is THE moment of Physical AI!
We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀
- Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions.
- It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.”
- Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks.
Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate.
The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community.
Welcome to the era of Physical AI.
HuggingFace: https://t.co/QW5h5pIWWM
Project Website: https://t.co/Jppa0gkn16
Code: https://t.co/aJgaLm5BaG
Introducing Cosmos 3: Our latest frontier model for Physical AI
Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation.
Today we’re releasing Super (32B) and Nano (8B) variants.
We are back. After one year of quiet building.
Introducing GENE-26.5, our first robotic brain that takes a major step toward human-level capability.
For years, robotics has struggled to learn from the world’s largest and valuable data source: Humans.
Solving it means rethinking the whole stack from the ground up:
- A robotics-native foundation model.
- A 1:1 human-like robotic hand.
- A noninvasive data collection glove for motion, force, and touch.
- A simulator that turns weeks of experiments into minutes.
GENE-26.5 is trained across language, vision, proprioception, tactile, and action. We designed a set of tasks to test how far we can go with this new paradigm.
Fully autonomous, 1x speed, one model, same weights. (Enjoy with sound on)
We are approaching the endgame for robotics.
And this is just a beginning.
Happy to see the release of DreamDojo! This is a great cross-team collaboration effort at NVIDIA. Real-time simulation is coming to our physical world 🦾🤖
Announcing DreamDojo: our open-source, interactive world model that takes robot motor controls and generates the future in pixels. No engine, no meshes, no hand-authored dynamics. It's Simulation 2.0. Time for robotics to take the bitter lesson pill.
Real-world robot learning is bottlenecked by time, wear, safety, and resets. If we want Physical AI to move at pretraining speed, we need a simulator that adapts to pretraining scale with as little human engineering as possible.
Our key insights: (1) human egocentric videos are a scalable source of first-person physics; (2) latent actions make them "robot-readable" across different hardware; (3) real-time inference unlocks live teleop, policy eval, and test-time planning *inside* a dream.
We pre-train on 44K hours of human videos: cheap, abundant, and collected with zero robot-in-the-loop. Humans have already explored the combinatorics: we grasp, pour, fold, assemble, fail, retry—across cluttered scenes, shifting viewpoints, changing light, and hour-long task chains—at a scale no robot fleet could match. The missing piece: these videos have no action labels. So we introduce latent actions: a unified representation inferred directly from videos that captures "what changed between world states" without knowing the underlying hardware. This lets us train on any first-person video as if it came with motor commands attached.
As a result, DreamDojo generalizes zero-shot to objects and environments never seen in any robot training set, because humans saw them first.
Next, we post-train onto each robot to fit its specific hardware. Think of it as separating "how the world looks and behaves" from "how this particular robot actuates." The base model follows the general physical rules, then "snaps onto" the robot's unique mechanics. It's kind of like loading a new character and scene assets into Unreal Engine, but done through gradient descent and generalizes far beyond the post-training dataset.
A world simulator is only useful if it runs fast enough to close the loop. We train a real-time version of DreamDojo that runs at 10 FPS, stable for over a minute of continuous rollout. This unlocks exciting possibilities:
- Live teleoperation *inside* a dream. Connect a VR controller, stream actions into DreamDojo, and teleop a virtual robot in real time. We demo this on Unitree G1 with a PICO headset and one RTX 5090.
- Policy evaluation. You can benchmark a policy checkpoint in DreamDojo instead of the real world. The simulated success rates strongly correlate with real-world results - accurate enough to rank checkpoints without burning a single motor.
- Model-based planning. Sample multiple action proposals → simulate them all in parallel → pick the best future. Gains +17% real-world success out of the box on a fruit packing task.
We open-source everything!! Weights, code, post-training dataset, eval set, and whitepaper with tons of details to reproduce. DreamDojo is based on NVIDIA Cosmos, which is open-weight too.
2026 is the year of World Models for physical AI. We want you to build with us. Happy scaling!
Links in thread:
Persistent world models are core to robotics 🦾 simulation 🖥️ next-gen gaming 🎮. We take an important step toward this goal.
Really awesome work by @lemonaddie0909!
w/ Shitao Tang, Min Shi, @AlvinLiu27, @JinweiGu98, @liu_mingyu, @lindahua.
Website: https://t.co/TmxBfSd3DB
PlenopticDreamer 🤩 we can now generate and re-render persistent 4D dynamic scenes from any perspective! This capability is critical for true video world models 🌎.
It's not just video generation — we are modeling the space of plenoptic functions 🎥✨.
https://t.co/2bqhWK8CrQ
Video generation, but 4D, dynamic, scene-consistent, and very long at the same time?!
Introducing 𝐏𝐥𝐞𝐧𝐨𝐩𝐭𝐢𝐜𝐃𝐫𝐞𝐚𝐦𝐞𝐫, 𝐦𝐮𝐥𝐭𝐢-𝐯𝐢𝐞𝐰 𝐯𝐢𝐝𝐞𝐨 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐰𝐢𝐭𝐡 𝐥𝐨𝐧𝐠-𝐭𝐞𝐫𝐦 𝐬𝐩𝐚𝐭𝐢𝐨-𝐭𝐞𝐦𝐩𝐨𝐫𝐚𝐥 𝐦𝐞𝐦𝐨𝐫𝐲! The scaling secret is very simple: an autoregressive paradigm with minimal 3D inductive bias, aided with a spatially grounded memory retrieval mechanism.
🌐 Project page: https://t.co/2h2YQdAQez
🌐 Paper: https://t.co/aIqhokIsP7
"The “ChatGPT moment” for physical AI is nearly here. 👏
Cosmos open world foundation models understand reality to create and reason about rich synthetic worlds—so self-driving cars and robots can learn from endlessly varied, physics-aware scenarios.
From a single image, 3D scene, or simulator trace, Cosmos generates realistic video, brings edge cases to life, and helps autonomous systems train for the long tail.
Watch the #CES2026 demo now 🎥 https://t.co/3dbbnJZBpB
3Dfy anything from a single image!
Very thrilled to announce SAM 3D. From an input image, select any object you want, 3Dfy it!
Blog: https://t.co/wtQLAqXTzW
Demo: https://t.co/tt3YqJlnRB
After a year of team work, we're thrilled to introduce Depth Anything 3 (DA3)! 🚀
Aiming for human-like spatial perception, DA3 extends monocular depth estimation to any-view scenarios, including single images, multi-view images, and video.
In pursuit of minimal modeling, DA3 reveals two key insights:
💎 A plain transformer (e.g., vanilla DINO) is enough. No specialized architecture.
✨ A single depth-ray representation is enough. No complex 3D tasks.
Three series of models have been released: the main DA3 series, a monocular metric estimation series, and a monocular depth estimation series.
The core team members, aside from me: @HaotongLin, Sili Chen, Jun Hao Liew, @donydchen.
👇(1/n)
#DepthAnything3
Our tech report for NVIDIA Cosmos-Predict/Transfer 2.5 is out 🤖! All of our latest advancements on world simulation for physical AI with open models 🤗
📃 https://t.co/HUhrpxGAnq
💻 https://t.co/t5eoPfReC3
NVIDIA Cosmos open models made major progress.✨
✅ Cosmos Predict 2.5 unifies text, image, and video world generation into one model that creates longer and more coherent simulations with improved grounding and efficiency.
✅ Cosmos Transfer 2.5 introduces precise, spatially controlled world transformations that are 3.5× smaller, faster, and higher in fidelity than before.
Together, these models push the boundaries of physical AI, enabling robots and agents to learn, reason, and operate in dynamically simulated worlds.
Read the @HuggingFace blog.
🔗https://t.co/AW7pEGSH8F
#NVIDIAGTC
I told my parents that I’d like to drop out of my cs phd program at Stanford a few months back.
They didn’t let me, we’re asian :)
So I graduated and started @moonlake.
@sharonal_lee and I saw urgency, and opportunities. Excited to share that we raised a 28 million dollar seed round to build the future for simulations and games.
Grateful for the angels @naval, @goodfellow_ian@stevechen, @JeffDean@rauchg, @emerywells, @JaredLeto, @chrlaf, alongside many more, and the venture partners that we are fortunate to work with: @moislamvc, Shaun Johnson, @chrmanning, Artem Barsukov, Elvin Hao, @mercebent, @veelarco and William Freiberg. If you're a founder and you're not partnering with them, you're making a big mistake.
Check out what we're about 👇
[1/N] 🎥 We've made available a powerful spatial AI tool named ViPE: Video Pose Engine, to recover camera motion, intrinsics, and dense metric depth from casual videos!
Running at 3–5 FPS, ViPE handles cinematic shots, dashcams, and even 360° panoramas.
🔗 https://t.co/1mGDxwgYJt
We build Cosmos-Predict2 as a world foundation model for Physical AI builders — fully open and adaptable. Post-train it for specialized tasks or different output types.
Available in multiple sizes, resolutions, and frame rates.
📷 Watch the repo walkthrough https://t.co/7er9nuFx3T
⚒️ Visit https://t.co/uwKKDwuwT0 for more
#NVIDIACosmos #PhysicalAI
Cosmos-Predict2 is our latest open video foundation model for Physical AI!
https://t.co/tD2WSn2Obc
If you’re at #cvpr2025, I would also love to chat with you about world models!
🚀 Introducing Cosmos-Predict2!
Our most powerful open video foundation model for Physical AI. Cosmos-Predict2 significantly improves upon Predict1 in visual quality, prompt alignment, and motion dynamics—outperforming popular open-source video foundation models. It’s openly available and fully customizable.
🔧 We’re open-sourcing:
📦 Pretrained model weights: https://t.co/DqehhT1Bc4
🛠️ Full codebase for inference & training/fine-tuning (with examples & tutorials using Robo data):https://t.co/cXf7tH6W2P
🏁 PBench, our new physics-driven benchmark: https://t.co/X9Rc99e2EC
💪 Our models are constantly improving. Stay tuned for even better ones!
learn more: https://t.co/bnJZ4xCax7
@NVIDIAAI #Cosmos
You think camera pose estimation is solved? How about
@nvidia's Racer RTX? https://t.co/VGfkOgnf0I
We challenge you with ⚡️ Lightspeed benchmark (as part of DynPose-100K release) — crazy wild videos with 𝗴𝗿𝗼𝘂𝗻𝗱-𝘁𝗿𝘂𝘁𝗵 camera poses!
https://t.co/KsKPfqVt4W
#CVPR2025
Excited to share ☀️Lightspeed⚡, a photorealistic, synthetic dataset with ground truth pose used for benchmarking alongside DynPose-100K!
Now available for download: https://t.co/iYr6wj53Eb
Paper accepted to #CVPR2025: https://t.co/Vt8eX9huGB
DynPose-100K will be fundamental to building next-generation models that understand and control the 3D visual world. Awesome execution from @_crockwell!
w/ @jt_tung@TsungYiLinCV@liu_mingyu David Fouhey.
Paper: https://t.co/eo22gncx9f
Website: https://t.co/h4IIIEfaD2
Cameras are key to modeling our dynamic 3D visual world. Can we unlock the 𝘥𝘺𝘯𝘢𝘮𝘪𝘤 3𝘋 𝘐𝘯𝘵𝘦𝘳𝘯𝘦𝘵?! 🌎
📸 𝗗𝘆𝗻𝗣𝗼𝘀𝗲-𝟭𝟬𝟬𝗞 is our answer! @_crockwell has curated Internet-scale videos with camera pose annotations for you 🤩
Download: https://t.co/KsKPfqW0Uu
Ever wish YouTube had 3D labels?
🚀Introducing🎥DynPose-100K🎥, an Internet-scale collection of diverse videos annotated with camera pose!
Applications include camera-controlled video generation🤩and learned dynamic pose estimation😯
Download: https://t.co/iL3iqqzYL8
@NVIDIA Cosmos has been recognized as both Best AI and Best overall of #CES2025!! Very honored to be part of the huge team effort. Cosmos will accelerate and shape the future of #PhysicalAI development!
#NVIDIACosmos world foundation model platform was voted Best #AI and Best Overall at #CES2025. 🎉
Read to see how Cosmos is accelerating #physicalAI development and why it's being recognized for its ability to "enable innovators at CES for years to come." 👇
🌌 NVIDIA Cosmos -- our World Foundation Model platform! Super excited to have made core contributions in multiple aspects. Physical AI is key to modeling the universe of worlds 🌎!
75-page tech report 📄: https://t.co/eUB816Fhzz
Try them out now 😲! https://t.co/7wKY3Tm006
Whether you're a researcher or developer, #NVIDIACosmos world foundation models are now openly available under our permissive license to the physical AI community via NGC & @huggingface. 🤗 #CES2025
See how Cosmos is democratizing #physicalAI development: https://t.co/CLbN2HGOT7