THE SMARTEST MODELS ON EARTH FIND A CHAIR AND STILL FAIL TO SIT ON IT 83% OF THE TIME.
Humanclaw dropped nine vision-language models into a simulated humanoid body for 1,218 episodes across scanned real houses. find the object, walk to it, sit on it.
Gemini 3.1 led the field. found the target 64.9% of the time, reached it 42.5%, sat down 16.8%.
Seeing was never the problem. legs and feet collide on 28 to 45% of steps because the model cannot tell where its own body is. The paper calls it a ghost, "fluent about the world but with no sense of its own body."
The clip is deepmind's dedicated robotics model, a different system and a good one. The demos come from specialists. The general models everyone calls almost-agi still can't find their own knees.