SpaceXAI's agentic Grok models and Robocurve's independent real-world robot benchmarks could pair perfectly: open evals of Grok driving arms or humanoids on data-center construction, satellite assembly, and factory tasks. Add space-specific suites for zero-g ops or lunar builds via Inspect Robots. Transparent physical AI progress for multiplanetary scale.
Introducing @robocurve, a Public Benefit Corporation to measure and report frontier robotics capabilities.
We build real-world evaluations for robots and publish results as a neutral third party.
LLM token output speed increases by 2-7x per year, with Fable-class models doubling every month.
If trends continue, LLMs could meaningfully control robots in real time by end of the year, or by 2029 at the latest.
Gemini 3.7 Flash just saturated one of our physical tool-use benchmarks at 92%. Gemini 3.6 Flash, released three weeks earlier, scored only 32%.
This marks a step change in the robotics capabilities of LLMs. 🧵
Robots can bake bread, cut carrots, and do karate kicks in demo videos. But how capable are they, really?
@robocurve measures how good AI and robots are in the physical world, from making a sandwich to building data centers.