"Does S1 exhibit physical prompt steerability: different prompts induce distinct behavior in the same environment?"
Yes! Watch S1 follow 4 different video prompt recipes in the same kitchen.
The first one is something far out of distribution -- "putting a plate in a toaster".
Thrilled to share what I've been working on!
This figure from the blog is what makes me excited about pre-training in robotics today. With scale and diversity, performance of in-context learners like S1 predictably improves and far exceeds conventional VLAs on OOD tasks! (Robotics is far more data-bottlenecked than LLMs and most real world use-cases are OOD)
Here's one of the most interesting interactions I've had with S1. We bought a skateboard to try and see if S1 could perform wheel assembly - a precision task that wasn't seen in pre-training. Prompted with a single video demonstration consisting of wrist camera streams, S1 consistently aligns the wheel with the axle. This kind of skill would have taken a couple hours of data + fine-tuning a few months ago. Even better - when it misses the alignment - it recovers from this mistake by re-grasping and adjusting the wheel - a sign of how S1 understands intent even in OOD scenarios!
The first time S1 flipped a pancake, we assumed pancake flipping must have been in its pre-training data.
We searched our whole pre-training data and found no examples of flipping.
S1 inferred the out-of-distribution task from one video prompt.
first signs of life late april.
few months later, 10+ min unseen tasks from one demo.
the interesting part is compositional generalization context isn’t just specifying the task, it’s letting the policy recombine behaviors learned during pretraining.
icl is getting interesting.
Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning: