Agentic navigation on large scale databases is tricky and I’d love to explain the challenges we met to you all. Besides, I’m also open to chat about building efficient and performant search agents (advertising our new LongHaness benchmark https://t.co/ONYSCxp8NT 😛)
GPT-6 Astra is extremely good at creating digital twins 🤯!
I fed it a few real-world episodes and it was able to create a near exact replica in Mujoco + Blender
One step closer to closing the real2sim gap
Introducing FLUX 3 Action.
An open weights 7B World Action Model that achieves first place on the RoboLab benchmark.
It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.
FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA.
Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson.
Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next.
FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together.
We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
CliffCompaction, our autocompaction tool for coding agents, is out!
Sessions run for millions of tokens, agents stay on track, and costs drop by up to 50%.
Paper: https://t.co/OdPp6Mt1Gh
Code: https://t.co/39Mg0ZJZ9C
PyPI: https://t.co/yOSIP3KCUH
Blog: https://t.co/cZdN9ZbURf
@EliJalbout@NVIDIARobotics Hi Elie, would love to talk more about this. Your DMs seem to be closed so I sent you a connection request on LinkedIn instead. Hope to chat soon!
https://t.co/1iTca1IILx
My first robotics paper is out 🤖
World models are basically robotics' smartest friend who thinks 7 times before answering any question. great with text, terrible at parties 🥺
So we asked: can we keep the brain and lose the lag? THAW-VLA 🧵
(1/n)
VLAs? WAMs? Are language and video models the right foundations for robotics?
Introducing Grounded Action Model (GAM): a new paradigm that builds robot foundation models on top of a pretrained 3D grounding model.
Ground first. Then learn to act. 🧵👇
Check it out, codebase and checkpoints are all released🥳
🌐 https://t.co/1iTca1IILx
📄 https://t.co/iVIZL1ta4T
💻 https://t.co/5G21Q3Vnlh
🤗 https://t.co/KfCwCQY9H7
(n/n)
biggest thanks to my advisor @yong_jae_lee for trusting me to set up our robot lab from scratch. it started as boxes of parts and a lot of googling. now it's real robots doing real tasks, and this paper 🤖 and to @_jadenpark for building it alongside me 🙏
(6/n)
so grateful to KIMLAB at UIUC (Sankalp Yamsani and Prof. Joohyung Kim) for hosting me this summer. learned a ton, got to see a lot of robots, and had a great time doing it 🙏
(5/n)
we tried our best to break it. swapped the student size, the backbone, the alignment layer, even the teacher (3 different world models). gain showed up every time. it's a real prior, not two networks that happened to vibe🫡
(4/n)
not gonna pretend this is a big-brain paper. it's very simple. but it hits 97.9% on LIBERO (98.8% on a good day), also improves on RoboCasa-GR1, and works on real single-arm AND bimanual robots.
Ockham's razor is the way 🪒
(3/n)
the recipe: take a frozen world model, run it over your training data, then add an additional cosine loss so a small VLA learns to match its features.
teacher stays frozen. student thaws the knowledge out. hence THAW 🧊➡️💧
at inference: 0.8B, 32 ms, 1.86 GB on a 5090⚡️
(2/n)
Today, we're introducing SimFoundry, our real2sim2real framework at NVIDIA GEAR that automatically turns real-world scenes into simulation-ready worlds from a single image or video.
Website: https://t.co/JB3kf3GlYm
Paper: https://t.co/pVhE1qWXtU
This work marks a major step for our team toward leveraging simulations and synthetic data for foundation model training and systematic policy evaluation at scale. Code will be open-sourced soon. Stay tuned!
🔥Our paper “Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents” was accepted to ICML 2026!
We show that post-training LLMs with RL to play a cooperative game teach generalizable cooperative reasoning.
1/n
We are grateful to all of the 17,491 reviewers who helped make #CVPR2026 possible. We are especially pleased to recognize the following Outstanding Reviewers, whose high-quality reviews (as judged by their Area Chairs) placed them among the top 5% of reviewers.
Check out Ariticraft 🦾 - a highly efficient agentic system that generates articulated 3D assets fully automatically at scale!
🚀 https://t.co/anSM87Li49