I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀
Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
Autoresearch just left the sandbox and entered the embodied world.
We are excited to introduce 𝐄𝐍𝐏𝐈𝐑𝐄: a system that drops frontier coding agents onto a fleet of real robots and hands them the entire loop:
reset the environment → search the literature → implement ideas and build the infra → train and deploy → self-verify → analyze the logs and rewrite the code → repeat, until the policy is reliable in the real world. No human in the loop.
Guided only by the robot's self-proposed, heuristic-based success signal, the agents hill-climb to 99% on dexterous real-world tasks: organizing pins into a box, seating GPUs, tying zip-ties.
We envision the bottleneck in robotics shifting — from building smarter algorithms to building the closed physical feedback loops an agent can finally turn on its own.
🔗 https://t.co/3tL2ArGo3v
From @NVIDIA@CMU_Robotics@Berkeley_AI
🧵
TopoRetarget enables a dexterous manipulation RL mimic pipeline from human demonstrations to real-world robot execution.
With TopoRetarget + RL tracking, we achieve zero-shot sim-to-real pen spinning on the Wuji Hand.🎉🎉🎉
We collect human hand-object motion with Wuji Glove, retarget the human demonstration to wuji hand while preserving hand-object interaction information, and train a RL policy to track the resulting hand-object reference.
Paper: https://t.co/e9TrpkFMpd
Project page: https://t.co/U0KncaOIaR
We release Wuji MJLab, an open-source MuJoCo environment for dexterous hand manipulation.
It includes a cube reorientation task, a sim2real pipeline, and the setup needed to reproduce the system.
We will be at ICRA booth 121 with a live demo and welcome discussions on dexterous manipulation.
Code: https://t.co/EpJMcozjnV
Contributors: Jielin Wu, @yaoshzh, @BergerHunger, Li Chengmeng.
This project is based on @kevin_zakka’s mjlab project.
No wonder they say teleoperation is dead
Peking University’s DAGroup released HumanNet, a massive 1 million hour dataset of human-centric videos that turns everyday internet footage into gold for embodied AI.
It’s got first-person and third-person views, super detailed actions, object interactions, tool use, and long sequences of real behavior. Everything’s cleaned up and annotated with 3D poses, SLAM tracks, and robot-ready labels.
The crazy part is,Training on just 1,000 hours of their first-person videos performs as well as (or even slightly better than) 100 hours of actual robot teleop data on downstream tasks.
Basically, human videos can now stand in for expensive robot data. Scaling laws for robotics just got a whole lot more realistic.
Human priors + smart curation = the future of world models.
🤖Low-data post-training can teach a VLA policy a new robot skill. But it also makes it too attached to the training demos.
We call this lock-in🔒: the policy can execute the post-training task, yet fails to respond to seemingly obvious prompt changes.
DeLock preserves steerability using only the policy’s own pretrained knowledge. No extra supervision needed!🚀🚀🚀
#Robotics #AI #EmbodiedAI #VLA
Excited to introduce Locomotion Beyond Feet!! A whole-body locomotion system that enables humanoid robots to crawl, climb, and recover using hands, knees, elbows, and torso, extending locomotion beyond legs alone.
Project Page: https://t.co/catSmExed6
1/7
‼️VLMs/MLLMs do NOT yet understand the physical world from videos‼️
In our recent work, we found that even the most advanced AI models still lag behind humans in one key aspect: reasoning about the kinematic properties of objects from videos.
Takeaways:
1. ChatGPT 5.1 leads overall among 21 advanced VLMs, followed by Gemini 2.5 Pro/Flash.
2. Grok 4.1 delivers impressive performance at the lowest API cost.
3. Qwen3-VL is the top-performing open-source model.
Read here: https://t.co/5lagvLNE37
🧵1/N
We tried to rethink modern robot learning policy design in this paper. TL;DR: Generative policies perform well NOT because of their distributional-learning formulation or their ability to capture “multi-modality.” Our proposed Minimal Iterative Policy (MIP) achieves similar performance to flow with much lower inference & training cost.
Some interesting takeaways:
1. “Multi-modal behavior” is a very nuanced concept in robotics. Robot learning cares about the conditional distribution P(a|o), where a is the action (chunk) and o is the observation.
The objective of generative modeling in vision/language is fundamentally different from the goal in control. In vision/language, we want diverse samples from the data distribution. In control, any action that leads to better downstream performance is sufficient.
Also, robotics has very little (labeled) action data, and they lie on a low-dimensional manifold. As a result, even if the marginal P(a) is multi-modal in many tasks (e.g., PushT), the conditional P(a|o) is often not multi-modal in sparse-data regions.
Empirically, we observe that the common explanation—“flow/diffusion policies win because demonstrations are multi-modal”—does not hold for most studied behavior cloning benchmarks.
2. Architecture and action representation matter more than flow vs. regression. We find that the strong performance of flow/diffusion policies is largely due to their modern architectures and action representations. UNet / Transformer designs and action chunking play a vital role. In fact, L1/L2 regression policies can match flow/diffusion when paired with the right architectures and chunking, except in a few high-precision tasks.
3. What actually helps in flow/diffusion? Stochasticity injection and supervised iterative computation.
To isolate these factors, we performed a “surgery” on flow to design Minimal Iterative Policy (MIP), a deterministic two-step policy that keeps only these mechanisms.
Surprisingly, with far lower compute, MIP is comparable to—or even better than—flow in most studied behavior cloning benchmarks.
Why? We hypothesize that these mechanisms provide good inductive biases that improve manifold adherence, especially for OOD states. Flow and MIP exhibit significantly better manifold adherence than regression.
I don’t think these conclusions are complete—there are still many mysteries. But I strongly feel we need a new subfield: “the physics/science of robot learning,” aimed at understanding fundamental learning mechanisms and principles for data-driven robotics.
As @ilyasut said for LLMs: “We’re moving from the age of scaling to the age of research.” For robot learning, I believe we must embrace the age of research to unlock better scaling to ignite the age of scaling.
Understanding Multi-View Transformers
Michal Stary @jgaubil@_atewari@vincesitzmann
tl;dr: DUSt3R self-attention is it secretly a diffusion model, and cross-attention is matching.
https://t.co/UR9agpjD8M
🚀 Agility Meets Stability (AMS) — one unified policy for humanoids that can dance, run, and balance like Ip Man 🥋.
By learning from heterogeneous data (human MoCap + synthetic balance motions), AMS achieves both dynamic agility and extreme stability in a single controller.
👉 https://t.co/pDsShPjdpV
Meet BFM-Zero: A Promptable Humanoid Behavioral Foundation Model w/ Unsupervised RL👉 https://t.co/3VdyRWgOqb
🧩ONE latent space for ALL tasks
⚡Zero-shot goal reaching, tracking, and reward optimization (any reward at test time), from ONE policy
🤖Natural recovery & transition
Excited to introduce TWIST2, our next-generation humanoid data collection system. TWIST2 is portable (use anywhere, no MoCap), scalable (100+ demos in 15 mins), and holistic (unlock major whole-body human skills).
Fully open-sourced:
https://t.co/fAlyD77DEt
VLA is great and all, but how do we further improve an existing VLA with more knowledge and learn from their own mistakes?
In this project led by @_wenlixiao, we mix online RL magic into VLA finetuning and let VLA improve themselves!
The end result? GPU assembling nonstop for hours!
Introducing RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning. https://t.co/FUqEQ26mzi
7 real robot tasks, 900/900 successes. Up to 250 consecutive trials in one task, running 2 hours nonstop without failure.
High success rate against physical disturbances, zero-shot, and few-shot adaptation
Our first step toward a deployable robot learning system.
Humanoid motion tracking performance is greatly determined by retargeting quality!
Introducing 𝗢𝗺𝗻𝗶𝗥𝗲𝘁𝗮𝗿𝗴𝗲𝘁🎯, generating high-quality interaction-preserving data from human motions for learning complex humanoid skills with 𝗺𝗶𝗻𝗶𝗺𝗮𝗹 RL:
- 5 rewards,
- 4 DR terms,
- Proprio. ONLY,
- NO history/curriculum.
Ready for agile, human-like 🤖? (Best with 🎧)
🔗 https://t.co/rT2CRb9msm 🎥
1/9