GPT-6 Astra generates the best zero-shot Real2Sim assets tested against Opus 5.5 and Fable 5.1. It’s also the most expensive.
We gave frontier models ~20 iPhone captures of real-world objects and asked them to recreate them as realistic, functional NVIDIA Isaac Sim assets:
🧵
Find the prompts and downloadable assets here: https://t.co/m6nRx6YsuD
We’re expanding our robotics evaluation work with a small number of partners. Tell us about your project: https://t.co/JPjd5dCkXN
All tests were zero-shot without human intervention. Inputs were 12-25 ordinary iPhone photos from different angles, plus 1-4 short videos of the objects and their mechanics.
We then placed, tested and rendered the resulting assets in our own Isaac Sim environment.
Even where zero-shot generation falls short, frontier models have become remarkably useful for asset generation. With a few days of iteration, we built a realistic, functional vacuum cleaner for use in our private evals.
Astra is dominating existing robot policies. But what can Astra do that existing policies could not?
We analyzed 10 tasks, 5 Episodes each on Astra, Pi0.5, Nvidia’s Cosmos3 Nano and Gr00t 1.7 on RoboLab.
1) Astra is by far the most reliable model in terms of object recognition:
GPT-6 Astra is the most significant leap in robotics I’ve seen in the past few years. It cracked RoboLab with a near-perfect score. Solid infrastructure + scaling ultimately outperformed the heuristics explored in small-scale studies. We’re definitely on the brink of physical RSI.
Source: https://t.co/JvQcTZAy5i
@tonyzzhao Exaggerating results causes a bad equilibrium for the entire industry.
Everyone else also needs to make up improvements as well to not be perceived as falling behind.