Let’s put them side by side. 👀
Right: A month ago, GPT-5.6 Sol barely managed to complete the two-disk task—even with a whole toolbox of helpers.
Left: GPT-6 made it halfway through the five-disk task, using only rendered camera images and robot joint angles as observations.
Why only halfway? You’ll have to ask Astra why it’s so hungry for tokens. 😂
We’re releasing a repo soon with our latest experiments on LLMs + harnesses directly controlling robot arms and dexterous hand across several simulation environments.
Once it’s public, bring your own LLM and harness and give it a spin! We’d love to see what you build—and how far you get before the tokens run out. 🤖
Stay tuned, and happy experimenting!
If Tibo could help me get some reset, the release would be even faster! @thsottiaux
Let’s put them side by side. 👀
Right: A month ago, GPT-5.6 Sol barely managed to complete the two-disk task—even with a whole toolbox of helpers.
Left: GPT-6 made it halfway through the five-disk task, using only rendered camera images and robot joint angles as observations.
Why only halfway? You’ll have to ask Astra why it’s so hungry for tokens. 😂
We’re releasing a repo soon with our latest experiments on LLMs + harnesses directly controlling robot arms and dexterous hand across several simulation environments.
Once it’s public, bring your own LLM and harness and give it a spin! We’d love to see what you build—and how far you get before the tokens run out. 🤖
Stay tuned, and happy experimenting!
If Tibo could help me get some reset, the release would be even faster! @thsottiaux
The recent attempt in agentic robotics has inspired me.
1. The grasping task usually has a good effect, because the grasping itself is closer to the processing of static scenes (including dexterous hand grasping!)
2. Collisions and avoiding collisions are difficult for LLM, and it is difficult to solve by yourself without human help (humans have some mobile experience to prevent collisions)
3. At present, dexterity operation is still the territory of imitation learning and reinforcement learning, but who knows how long it will last?
Final question, in addition to collecting data, what skills do you currently master within five years cannot be replaced by LLM?
Thanks Pei Zhou for the interesting demo. Thanks @alex_kai2020@Lixuan_thu for communicating and helping.
@YiMaTweets Have studied this paper in detail. Representing camera pose as a dense 3D-to-2D map instead of a traditional matrix is a great simplification. Super elegant way to unify shape and pose generation!
@liuziwei7@_akhaliq Soooo great ! Of great insights to move beyond just connecting separate vision and language modules , and it can get us back to first principles for how vision and language should be integrated
@xwang_lk Exactly. There's a long way for us to explore where human intelligence really lies in (Even for that "ez"question) and how to represent them in machine
@YiMaTweets Yeah It's the same principle that underpins almost everything from probabilistic robotics to the iterative state updates in recurrent networks and Transformers.