Can we build a robot policy foundation model🤖 that incorporates 3D observations?
🚀Introducing FP3, a first 3D foundation policy for robotic manipulation. Using only 80 demonstrations, FP3 can solve a new task with over 90% success rates in the wild!
Let's dive in! 🧵1/N
Human demos contain rich hand-object skills. The challenge is the embodiment gap. Human and robot hands differ in shape, joints, and contacts.
Introducing ConTrack, our latest work that turns contact-rich human hand-object demos into robot motions.
#CVPR2026 GR3D 🧊 — A single VLM that grounds in 2D, grounds in 3D, and reasons with visual chain-of-thought — all at once.
Excited to share our paper, Grounded 3D-Aware Spatial Vision-Language Modeling!
Excited to announce that FP3 has been selected as an ICRA 2026 Best Paper Award Finalist on Robot Learning! 🎉
Unfortunately, I won’t be able to attend in person in Vienna, but our advisor @gao_young will be presenting the work today. Please stop by if you are at ICRA!
📍 Session: TuAT1.3
🕦 Time: June 2 11:20–11:30
Glad to share that our paper FP3 has been accepted to #ICRA2026! 🎉Huge thanks to my co-authors and collaborators, it's been a long ride.
In FP3, we demonstrate that 3D perception is a strong ingredient for learning robot foundation models. It’s been about 1 year since we wrapped up this work, and during this year, video world models have advanced insanely fast to push the vision community into a “do we still need explicit 3D?” debate. The same question also applies to robotics. Classic stacks rely on rigorous geometry, but will end-to-end learning make implicit 3D sufficient?
Excited to explore the real boundary.
VLA/VAs are doing well on short skills like pick-and-place. But real tasks rarely stop after one action, they require 1) many interdependent steps, 2) progress tracking, and 3) recovery from mistakes.
In our paper LoHo-Manip, we address long-horizon manipulation with trace-conditioned VLA planning: a task manager tracks what’s done, plans what remains, and guides execution with visual traces.
Introducing EgoVerse: an ecosystem for robot learning from egocentric human data.
Built and tested by 4 research labs + 3 industry partners, EgoVerse enables both science and scaling
1300+ hrs, 240 scenes, 2000+ tasks, and growing
Dataset design, findings, and ecosystem 🧵
Ever want to have a single policy to control diverse robots as well as different dexterous hands, or to observe the emergent behavior under cross embodiment training?
Introducing our #CVPR2026 paper XL-VLA, Cross-Hand Latent Representation for Vision-Language-Action Models.
Learning from vast video data allows point-flow-based planners to create generalizable task plans, guiding robots via future keypoint trajectories.
But how can we ensure low-level execution doesn't become the system's bottleneck?
Introducing our #ICLR2026 paper, HinFlow: Translating Flow to Policy via Hindsight Online Imitation
Turning UMI-recorded end-effector trajectories into manipulation-ready humanoid whole-body actions is definitely challenging. Congrats to @RuiqianNai and team for delivering such impressive results!
🤖 Can we demonstrate humanoid complex whole-body manipulation skills without a physical robot present?
Introducing HuMI: A portable, robot-free interface for learning diverse humanoid manipulation tasks.
📄 https://t.co/bKdgzgTnQb
🌐 https://t.co/K7nWFORg3b
"Cross-embodiment" is a sign of generalization. We’ve seen huge progress in manipulation and navigation — but what about humanoid whole-body control? Can ONE policy control multiple different humanoids?
Meet our #ICRA2026 work 🦅EAGLE: Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control.
Instead of brute-force URDF / morphology domain randomization, we iteratively distill specialists into one generalist. We also find that embodiment-aware representations matter for policy learning.
🔗 website: https://t.co/ox6xNcu5zz
📜 arXiv: https://t.co/ddLZi9smkM
Can we bridge the Sim-to-Real gap in complex manipulation without explicit system ID? 🤖
Presenting Contact-Aware Neural Dynamics — a diffusion-based framework that grounds simulation with real-world touch.
Implicit Alignment: No tedious parameter tuning.
Tactile-Driven: Captures non-smooth contact events.
Consistent: Stable predictions in contact-rich tasks.
𝑪𝒐-𝒕𝒓𝒂𝒊𝒏𝒊𝒏𝒈 is a promising way to scale Large Behavior Models (LBMs) beyond robot data, yet the data and training recipe are far from settled. 🤔
We present a large-scale empirical study leveraging 4,000h of robot/human data and 50M vision-language samples, evaluating 89 policies across 58,000 simulation rollouts and 2,835 real-world trials. 🤖📊
https://t.co/jMlZWdXexl
Work done during my internship at @ToyotaResearch.
Glad to share that our paper FP3 has been accepted to #ICRA2026! 🎉Huge thanks to my co-authors and collaborators, it's been a long ride.
In FP3, we demonstrate that 3D perception is a strong ingredient for learning robot foundation models. It’s been about 1 year since we wrapped up this work, and during this year, video world models have advanced insanely fast to push the vision community into a “do we still need explicit 3D?” debate. The same question also applies to robotics. Classic stacks rely on rigorous geometry, but will end-to-end learning make implicit 3D sufficient?
Excited to explore the real boundary.
Can we build a robot policy foundation model🤖 that incorporates 3D observations?
🚀Introducing FP3, a first 3D foundation policy for robotic manipulation. Using only 80 demonstrations, FP3 can solve a new task with over 90% success rates in the wild!
Let's dive in! 🧵1/N
VLA models are booming—but one fundamental question is yet seldom answered:
How does the choice of the base VLM affect VLA performance���️
We run a large-scale systematic study to answer this, in collaboration with @Alibaba_Qwen
https://t.co/sfCyqfID50
How far can we push the limit of in-hand manipulation dexterity?
Introducing our work on motion capture: DexterCap & DexterHand !
DexterCap: A high-fidelity motion capture system for intricate in-hand manipulation motion.
DexterHand: A dataset featuring true in-hand dexterity, reorientation, finger gaiting, and even manipulating a Rubik's Cube like a speedcuber ! 🧩
- Project Page: https://t.co/9EASACaY6p
- Experience it via our online interactive visualization: https://t.co/IHCkbHBuuq
#Animation #CharacterAnimation #MotionCapture #Graphics #EmbodiedAI #DexterousManipulation
Thrilled to announce that Spirit v1.5 from Spirit AI has officially reached #1 on the RoboChallenge leaderboard, outperforming pi0.5! 🚀
Scaling end-to-end embodied AI is a collective journey. We’ve been deeply inspired by the community—their work continues to push the entire field forward. Therefore, we are open-sourcing Spirit v1.5 today. We can’t wait to see what the community builds with it. 🛠️
🏆 Leaderboard: https://t.co/AE4kYUUik1
💻 Blog: https://t.co/80u37S63fe
💻 Open Source:
Code: https://t.co/uj0kQr9tpl
Model: https://t.co/AS3kuI9DpG
#SpiritAI #EmbodiedAI #Robotics #OpenSource #Spiritv1.5
Meet ACE-F — a novel, foldable teleoperation platform for collecting high-quality robot demonstration data across robot embodiments.
Using a specialized soft-controller pipeline, we interpret end-effector positional deviations as virtual force signals to provide the user with force feedback, without requiring expensive sensors! ACE-F simplifies control for a diverse array of robot platforms, making complex tasks that require dexterous manipulation highly intuitive.
Check out the project website here!
https://t.co/gZJkNxsfg1
How do you teach a robot to do something it has never seen before? 🤖
With human data.
Our new Human0 model is co-trained on human and humanoid data. It allows the robot to understand a novel language command and execute it perfectly in the wild without prior practice.
Real-world success rate: ~100%.
Watch it happen 👇