11 demonstrations, then the robot taught itself the rest! 🎯
@Gobanorobotics, out of Nantes and London, published Toutatis v1 this month.
It's a reinforcement learning engine aimed at the part of robotics that doesn't make it into demo videos: getting from "it works" to "it works 999 times out of 1,000."
A 95% policy fails 50 times per 1,000 cycles. A strong lab result at 99% still fails 10 times. Production wants roughly one.
Human demonstrations can't close it. They teach a robot what a person did, not what maximises success rate or removes the rare failure that stops a system running unsupervised.
The engineering choice worth noting is where the RL actually runs. Gobano trains a task-specific world model from the same demonstrations, freezes the visual encoder, and caches the embeddings.
RL then operates on that compact latent state instead of backpropagating through the full vision stack on every update. Around 10x faster policy training than their own end-to-end visual RL setup.
→ Ball picking: 11 human demos, 24/100 on the initial imitation policy, 108/108 after 854 real-world rollouts
→ Zip-tie insertion: 107 demos, 3,375 rollouts, 31/31 on the final evaluation
→ Towel folding: 99/100
→ Simulation: 100% on Square and 99.8% on Tool Hang over trailing windows of 1,000 rollouts
→ Training runs on a gaming PC, and deployed policies run locally on a gaming laptop at roughly 3 ms per control step
They also built a separate "Reliability" training mode that optimises repeated success directly, since maximising average task reward lets rare failures hide behind routine wins.
Europe keeps showing up in physical AI 🇫🇷🇬🇧
🔗 Read the blog here: https://t.co/3AsfRP2B33
~~
♻️ Join the weekly robotics newsletter, and never miss any news → https://t.co/GoA3ZuwoPB
Demos are just the starting point.
Toutatis v1 is our RL engine for pushing robot policies from human demonstrations to 99+% performance & production reliability — in days.
From demos to deployment.
https://t.co/y8Gz0e9HcG