We just hit 100M ARR, 10 months after our first deployment.
One of the fastest growing physical companies in human history.
Deeply grateful to our team and partners for making this achievement possible. We're just getting started.
We just hit 100M ARR within 10 months of starting deployments.
We are in factory lines. On construction sites. In kitchens. In data centers.
Cleaning. Welding. Building. Cooking.
Deploying.
If S1 makes a mistake, it tries again.
For example, the input video prompt here attached the wheel just once. During deployment the wheel didn't align properly, so S1 re-aligned it.
By watching the prompt, S1 understands the goal and improvises to get there:
Thrilled to share what I've been working on!
This figure from the blog is what makes me excited about pre-training in robotics today. With scale and diversity, performance of in-context learners like S1 predictably improves and far exceeds conventional VLAs on OOD tasks! (Robotics is far more data-bottlenecked than LLMs and most real world use-cases are OOD)
Here's one of the most interesting interactions I've had with S1. We bought a skateboard to try and see if S1 could perform wheel assembly - a precision task that wasn't seen in pre-training. Prompted with a single video demonstration consisting of wrist camera streams, S1 consistently aligns the wheel with the axle. This kind of skill would have taken a couple hours of data + fine-tuning a few months ago. Even better - when it misses the alignment - it recovers from this mistake by re-grasping and adjusting the wheel - a sign of how S1 understands intent even in OOD scenarios!
Excited to see S1 finally go public! 🚀 This release also marks exactly one year since I transitioned from vision research to robotics. If I had to share the single biggest lesson from that journey, it’s this:
Before asking "How do we train a better model?", we need to ask "How should the model behave at inference time?"
In academia, we tend to attack problems the same way: set up a benchmark, tweak the architecture or training recipe, and watch the score climb. Today, much of robotics research runs this similar loop: pit VLAs against WAMs, introduce a new loss function, test a different action representation. Model performance on benchmarks keeps improving, but inference-time behavior stays simplistic and fixed.
But my past year at @SkildAI taught me that raw model performance isn't the ultimate driver of real-world success. Inference-time behavior is.
The real world guarantees high stochasticity and endless out-of-distribution scenarios. The dominant deployment paradigm (taking the past X seconds of observations to predict the next Y seconds of actions) won't get us to general-purpose robotics, because every robotic action has compounding consequences. A vision model can guess wrong on a static test set with zero impact on the next sample. A robot cannot. It operates in a closed loop: its own mistakes create the very distribution shifts it must then survive. "Collect more data, fine-tune, redeploy" is not a scalable answer. To truly generalize, a model must adapt on the fly.
So the order of operations has to invert. First, define an inference-time behavior scalable and robust enough to handle the real world. Then work backward to shape data collection and model training. Everything in the S1 release stems from this inversion.
Language modeling already taught us this lesson: LLMs didn't crack complex multi-step math just because someone threw more tokens at pre-training. Instead, the breakthrough came from changing the inference-time behavior first (reasoning step by step) and then redesigning training around it. Betting that data scale alone will magically produce emergent physical skills is just as naive. Robotics needs the same shift.
S1 is proof that working backward from real-world inference changes everything. And this is just the starting point.
Check out the release below, and enjoy the rest of your day! 👇
I am posting this as I am deploying S1 on a customer site. Words can't describe how excited I am.
This was a tremendous effort from the team across data, infra, training, inference, and hardware. Personally I think it is a giant step closer to general-purpose robots creating real value in our lives.
For homes, you can easily teach your robot a new task, or tell it your preferred way of doing a task.
For businesses, you can plug in your SOP and start seeing value creation from Day 1.
Now the foundation has been set, time to bring it to the real world. Still a lot more work to do, stay tuned.
Personally, working on the S1 release with the team has been one of the most satisfying accomplishments of this year. I’ve spent countless late nights discussing some of my favorite topics—like affordances and how we can build systems that truly generalize. We’ve celebrated every step change in generalization: from new object arrangements, to different objects, to where we are today—being able to perform completely unseen tasks. When I started working in robotics, I never imagined we’d reach a point where a robot could perform a task it had never been trained on, simply through prompting.
And we’re proud to say that no other research group has yet demonstrated generalization to completely unseen tasks, or to long-horizon tasks like making pancakes.
Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning:
In-context learning for robotics is here.
- Long-horizon tasks over 10 minutes long
- Never seen during pre-training
- Prompted with one video, no fine-tuning
We are building intelligence from the foundations up, not from the top down.
This is pretty unbelievable where robotics is getting to in the last 4 months. Nearly zero demos of anything like this before April.
Now feels like we are seeing something new every day. Sure, these are controlled environments and cherry picked, but the pace of advancement is clear.
a robot that knows 100 tasks still needs new training for the 101st. unless you can just show it the new one? :)
robots are finally starting to learn how to learn.
Skild is building that future with S1 and in-context prompting!!
Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning:
we missed our entire YC batch to build in Ukraine.
we went from idea to field testing, landed $500M in LOIs, and sold 45 units to a special forces unit
all in 8 weeks.
some things you just can’t build in SF
finally did a major twitter rebrand: deleted everything I posted in 2018
no more sad quotes in random languages or deep thoughts I tried to express before I even spoke english 😔
anyway, here are some highlights from my emo era: