We put GPT-6 Astra in the RoboDojo. 🥋🤖
The RoboDojo Team conducted a comprehensive evaluation of GPT-6 Astra as an embodied agent, including:
• RoboDojo Sim & Real, compared with GPT-5.5 and DeepSeek-Flash
• Humanoid high-level control
• Dexterous piano playing with RoboPianist 🎹
• A systematic study of in-context learning (ICL)
Our key takeaway:
GPT-6 Astra demonstrates remarkably strong semantic and spatial understanding, together with impressive in-context adaptation.
At the same time, physical commonsense remains a clear bottleneck — revealing an important gap between understanding the world and truly reasoning about its physics.
Full report & demos:
https://t.co/EwLGoPxWqr
@_wenbozhang (project lead), @wenhaocha1, @frankzydou, @JinWeiyang18434, @YutaoOuyang, @minifullcapsule, @x_h_ucb, @YutaoOuyang, @YueChen614
The only tutorial you should attend at #ECCV26 (kidding 😂).
But I am incredibly psyched to be doing this with the best bunch out there. I strongly feel this is a timely tutorial!
We will present general approaches alongside our learnings from (post)-training impactful models like Flux2/3.
Tutorial website: https://t.co/eig5dg3xwP
Save your calendars!
It's been a hot robot summer 🤖 Robots can now drive cars, fold laundry, and cable data centers. But the gap between 95% and 99.9% task success is the gap between a demo and a business. Our take on physical AI, and where the $40T market gets won https://t.co/YDZpdXZPzt
Action chunking — especially executing long action sequences open-loop — is widely used in imitation learning for robotic manipulation. Why is it so effective and do we really need it? We find a key reason:
Long open-loop execution helps short-context policies imitate non-Markovian experts.
With this insight, we show how to move beyond open-loop execution: extending policy context restores reactivity while achieving even higher task performance.
🧵(1/5)
AI research is becoming recursive. Quant research is next.
Introducing AQuA 🌊 — recursive self-improvement for quantitative research, powered by GPT-5.6 Sol.
Discover factors. Train models. Learn from experiments. Repeat.
https://t.co/x494WPGqXz
🎥👇
Robot learning needs human data.
Hands are where intent becomes action: grasping, tool use, bimanual coordination, object handoff, recovery after contact.
For learning from human interaction, accurate hand pose is not just an annotation detail. It is the bridge between pixels, objects, contact, and motion.
Here are 12 hand-object interaction datasets worth knowing:
#RobotLearning #ComputerVision #HandPose #HumanData #EmbodiedAI
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**.
Meet the Model Discovery Agent (MDA). 🧵
Super excited to share Dyna-2, the first robot foundation model trained on over 1 million hours of human data. At this scale, we saw the emergence of a cross-embodiment transfer scaling law: training on increasing amount of human data not only improves model prediction on held-out human data but also on robot data the model has never seen before 🤯
We validate that this transfer scaling law does translate to on-robot performance across 3 different robot platforms, and uncovers that both training objective and data matter greatly for this emergence. Crucially, our results position video as a new scaling axis for physical AI.
Beyond these scaling-law oriented results, we also dissect Dyna-2 greatly, going in-depth on its many capabilities, including language following via world modeling, enhanced robustness and precision, zero-shot customer site deployment, and some cool video generation results!
Read our technical blog post for more details! I do think this is a very important result that shows a very different path for robot foundation models from the ones we are marching on. Really excited for the road ahead. We have more releases coming, stay tuned!
We live in a multimodal world. We see, talk, act, and dream.
Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch?
Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition
So what’s the path forward?
Introducing: Towards Physics of Multimodal Pretraining (https://t.co/xgfbdDPFq4)
We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
today we're launching automatic captioning at Instance.
send us your robot data, and every episode is segmented into captioned subtasks, graded success/fail plus 1-5 on quality and speed.
if this could be useful for you, reach out and we'll caption an episode for free!
Introducing Waddle Labs: Claude Code for robots.
Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes.
@yiding_song@theWaddleLabs
I just released a first video deep dive into my LeHome Challenge solution (🥇 sim round, 🥈 real-robot final at ICRA 2026).
I already shared what I did and the video series is about why: the reasoning, the alternatives I rejected, and the intuition behind each choice.
I split the whole video explanation into 3 parts: Part 1 is about RL for VLAs 👇
We gave popular agents (Claude Code & Codex) a browser-based visual interface and let them command a robot via MCP tools — like a human with a mouse and keyboard — and it worked!
Introducing VIA: Visual Interface Agent for Robot Control. 🧵
We scaled robot policies to 8K timesteps of visuomotor context, orders of magnitude beyond current SoTAs, at constant inference latency.
Introducing RoboTTT 🤖
With minutes of experience in context, our robots:
🎥 one-shot imitate human video demos
📈 improve themselves during deployment
🛡️ recover from perturbations
⚙️ complete a 5-minute, 10-stage assembly end to end
🌐 https://t.co/4OoK0e9tKa
Dive in 🧵
We’re excited to share our work on cross-embodiment dexterous manipulation:
The Unified Hand Action Space (UHAS) — a new representation that allows a single policy to control robotic hands with completely different kinematic structures and numbers of fingers.
🎥 This video shows one policy simultaneously controlling four different hands.
Project: https://t.co/m37qyfYcpX
🧵 Thread
1/9 🤖
Super excited to share the last paper of my PhD: "Hallucination in World Models is Predictable and Preventable"✨
We train a 350M-param generative world model on a large dataset w/ 210 tasks and show that we can predict *when* hallucination happens and use that to fix it!
🧵1/n
This Friday, @caizhongang joins our Video Generation & Reasoning Journal Club to present "Can Models Think Without Language?"
His talk explores a new paradigm: using video, not text, as the substrate of model reasoning. Unlike language, video naturally captures rich spatial and temporal structure, opening new possibilities for how models perceive, reason about, and make predictions about the world.
Friday, Jun 26
7:30 PM PT · 10:30 PM ET · Sat 10:30 AM Beijing
Online via Zoom
RSVP: https://t.co/U8ZbPhK7p4
More info: https://t.co/4DBNEqutBW