"Cross-embodiment" is a sign of generalization. We’ve seen huge progress in manipulation and navigation — but what about humanoid whole-body control? Can ONE policy control multiple different humanoids?
Meet our #ICRA2026 work 🦅EAGLE: Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control.
Instead of brute-force URDF / morphology domain randomization, we iteratively distill specialists into one generalist. We also find that embodiment-aware representations matter for policy learning.
🔗 website: https://t.co/ox6xNcu5zz
📜 arXiv: https://t.co/ddLZi9smkM
Touch has become a key ingredient for dexterous robot manipulation, but which tactile sensor should you actually use?
We're excited to introduce TacO, a benchmark for evaluating tactile sensors across real-world manipulation tasks.
🌮 TacO Benchmark provides:
Cross-modality comparison of different sensors (vision-, magnetic-, acoustic-, and resistive-based tactile sensing)
Standardized imitation learning framework with all kinds of tactile sensors
Open-source code, data, and hardware
Our experiments show that there is no universally best tactile sensor😌(even the expensive ones). The right choice depends on the tasks. We hope TacO helps the community build and evaluate the next generation of tactile sensors and how to use them to improve vision-based robot policies.
RL training for LLMs involves exposure to problems in the “Goldilocks zone” of difficulty: not too hard, not too easy.
But difficulty is not the only thing that matters.
Problem type matters, too.
And LLMs do not see or organize problem types the way humans do.
This is the starting point of Manifold Bandits.
The fundamental issue is this: when using RL to train LLMs, there will always be problems that are more productive for the policy at a given training iteration. Across a dataset, there are an astronomical number of possible batch orderings; uniform sampling is unlikely to consistently land on the most useful ones by chance.
This is why many adaptive curriculum learning and difficulty-aware methods focus on the policy’s “optimal learning zone,” avoiding examples that are already trivial or still impossible for the LLM to solve.
But problem type is usually treated differently.
Task decompositions within a training set are often manually defined according to human semantics, or ignored entirely, with an entire dataset/domain treated as “one task,” despite the immense heterogeneity seen even in small training datasets.
It is generally understood that LLMs may not share our concept of difficulty. But it seems less recognized that LLMs may not share our concept of problem type, either. In both cases, training the LLM according to our own semantics can obscure fine-grained learning dynamics that could otherwise be exploited.
In this work, we seek to derive an adaptive curriculum learning method that caters not only to the policy’s ability, but also its perception: its latent organization of tasks.
Driven by the manifold hypothesis and the Bitter Lesson, we leverage computation to derive curriculum structure from the policy’s latent geometry, then use Bayesian inference over that structure to guide search and learning over the training set. The result is Bayesian Manifold Curriculum (BMC): an algorithm that does not just try to find problems of the right difficulty, but instead orchestrates training effort across diverse and interacting problem types.
More technically, we frame problem sampling for LLMs as a manifold-structured bandit problem with endogenous non-stationarity. The tl;dr is that the problems/arms are diverse, structured, and interact through policy updates, so the typical non-stationary bandit framing is not quite the right fit for training LLMs. BMC is derived from this new framing.
🌐Website: https://t.co/iKVjNK92up
📰Paper: https://t.co/j08GktD6Qv
💻Code: https://t.co/ki6O3ojMWO
🧵1/N
For folks working on WM/WAM…, you should definitely read this paper on world model hallucination.
Very solid work, and it’s fully open-sourced.
Congratulations to Dr. Nick on his graduation, all the best! We will miss you! 🙌
Super excited to share the last paper of my PhD: "Hallucination in World Models is Predictable and Preventable"✨
We train a 350M-param generative world model on a large dataset w/ 210 tasks and show that we can predict *when* hallucination happens and use that to fix it!
🧵1/n
If the end goal of robot hands is to perform human motion, then we should optimize the hardware design with human motion - and on a large scale!
We can generate both a high-dof generalist hand, and also low-dof specialized hands from human demonstration.
What if a robot could learn the physics of a soft object, how a towel folds or lift a plush toy, by watching you play with it from a first-person view?
Introducing EgoPhys. We build deformable twins from one egocentric RGB video using a generalizable material codebook.
[1/6] 🧵
Introducing ABC: open data, training, and infrastructure for robotics.
We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques.
@arthurallshire@Cinnabar233@adamrasb@redstone_hong@davidrmcall
In the era of 10T-parameter foundation models 😈, can we still build next-generation world representations using 🦥 a single RTX 3090 GPU?
We introduce WorldString (a name inspired by string theory), a tiny neural network representation for reconstructing dynamic world instances, ranging from articulated motion and skeletal skinning to soft-body deformation. 🔥
We view object representation as the core primitive of world modeling: the basic executable unit from which perception, interaction, simulation, and generation can be built. With a compact neural basis, WorldString aims to connect dynamic object modeling with neural simulation and foundation models, enabling a scalable path toward richer representations of the physical world.
#CVPR2026 GR3D 🧊 — A single VLM that grounds in 2D, grounds in 3D, and reasons with visual chain-of-thought — all at once.
Excited to share our paper, Grounded 3D-Aware Spatial Vision-Language Modeling!
For dexterous manipulation, I mean, beyond grasping, one goal is to strike a balance between task achievement (🎯task reward) and human likeness (✍️style reward).
How to do that? You might find an answer in ConTrack. We've open-sourced the code, check it out: https://t.co/NGTIsGVri4
btw, Yutong is presenting the work at CVPR 2026; folks in Denver can catch him and ask for more details! 🤭
#CVPR2026 #Robotics
Human demos contain rich hand-object skills. The challenge is the embodiment gap. Human and robot hands differ in shape, joints, and contacts.
Introducing ConTrack, our latest work that turns contact-rich human hand-object demos into robot motions.
Excited to share that ARI (Assured Robot Intelligence) is joining @Meta!
When we co-founded ARI a year ago, the mission was clear: build humanoid intelligence for the real world.
Joining Meta Superintelligence Labs (MSL), we'll continue advancing frontier robotics models toward physical superintelligence in the physical world.
Huge thanks to my co-founders, the incredible ARI team, and our investors led by @aixventureshq for backing this from day one.
This is just the beginning.
Excited to share that Assured Robot Intelligence (ARI) has joined @Meta to help build the future of humanoid intelligence!
When we started ARI one year ago, our mission was clear: achieve physical AGI. Through deep customer engagements and real-world deployments, it became clear to us that serving the massive opportunity ahead requires training a truly general-purpose physical agent.
We believe this agent will be humanoid — and that scaling will come from learning directly from human experience, not teleoperation alone. Meta’s ecosystem brings together the key components needed to make this vision possible. We will be joining Meta Superintelligence Labs (MSL) to help bring personal superintelligence into the physical world.
We are incredibly grateful to the brilliant minds, robotics researchers, engineers, partners, and supporters who have worked with us on this journey. Thank you to our investors and angels, led by @aixventureshq , for believing in our mission.
This is just the beginning.
VLA/VAs are doing well on short skills like pick-and-place. But real tasks rarely stop after one action, they require 1) many interdependent steps, 2) progress tracking, and 3) recovery from mistakes.
In our paper LoHo-Manip, we address long-horizon manipulation with trace-conditioned VLA planning: a task manager tracks what’s done, plans what remains, and guides execution with visual traces.
Reasoning VLAs can think. They just can't think fast. Until now.
Introducing FlashDrive⚡
🚀 716 ms → 159 ms on RTX PRO 6000 (up to 5.7×)
✅ Zero accuracy loss
FlashDrive = streaming inference + DFlash speculative reasoning + ParoQuant W4A8
Real-time reasoning for autonomous driving is here!
https://t.co/zWIBhyJ5QN
Introducing GEN-1.
Our latest milestone in scaling robot learning.
We believe it to be the first general-purpose AI model to master simple physical tasks.
99% success rates, 3x faster speeds, adapts in real time to unexpected scenarios, w/ only 1 hour of robot data.
More🧵👇
We’re building UWLab, a shared ecosystem for training robot policies in simulation and transferring them to the real world, built on Isaac Lab.
This includes the full OmniReset codebase, along with tasks, algorithms, and deployment in one clean, modular stack: https://t.co/PLX1fzPiSU
We’re releasing OmniReset, a framework for training robot policies using large-scale RL and diverse resets for contact-rich, dexterous manipulation.
OmniReset pushes the frontier of robustness and dexterity, without any reward engineering or demonstrations.
Try the policies yourself in our interactive simulator! https://t.co/3hW3nYx2vD
(1/N 🧵)
Introducing EgoVerse: an ecosystem for robot learning from egocentric human data.
Built and tested by 4 research labs + 3 industry partners, EgoVerse enables both science and scaling
1300+ hrs, 240 scenes, 2000+ tasks, and growing
Dataset design, findings, and ecosystem 🧵