This seems like a pretty significant advancement in in-context learning. The model learned out-of-distribution tasks it was never trained on, like making a pancake, after watching a single video demonstration of the task, with no fine tuning.
Quoting from the blog:
'We believe the main purpose of pre-training — if not the only one — is to enable, for robotics, the same shift we saw from BERT to GPT-3 (Brown et al., 2020): in-context learning, or the ability to learn immediately from one or a few examples.
One year ago, we released an in-context learner for locomotion that adapts by accumulating live experience in its prompt (Liu et al., 2025). Today we are sharing manipulation results from our flagship robotic foundation model S1, built from the ground up as an in-context learner. Show it a video of a task, short or long, seen or unseen, and it executes.'