Riyan and I are excited to announce that we got accepted into a16z @speedrun.
At Sonder, we're building your digital twin.
We all know AI is here, yet only a few of us have really changed our workflows around it. I've tried and it is not easy. There's friction at every step & it still hasn't learnt to be me.
As the world learns to work with AI, your digital twin will make your experience smoother than mine.
We're hiring, reach out if you or someone you know is obsessed with on-device inference, teaching models how people work, or shaping how humans and AI work together.
You had to learn to use a computer. What if it learned to be you?
Introducing Sonder Labs, an applied AI lab building your digital twin.
Today, we’re excited to share that a16z @speedrun is backing us.
Our work starts with vision models small enough to live privately on your laptop - always on, seeing what you see. The more your twin learns about your workflows, priorities, and relationships, the less you have to explain.
And when it sees a way to help, it prompts you before you prompt it.
Ajit and I met deploying some of healthcare’s earliest LLM systems at Anterior. Working alongside nurses, we saw early what the rest of knowledge work is now confronting. There’s already more intelligence than we know how to use. We’re the bottleneck.
The tools we use capture what we produce, but lose much of what makes it ours: the iterations and rejected options behind the final result, the taste that defines what good looks like, and the intent behind our decisions. And the work doesn’t stay still. Priorities shift, workflows evolve, and exceptions become routine.
Your world can’t be described once. It has to be learned continuously.
Generic models are powerful, but they can pull work toward the average. We’re betting on AI that helps people become more themselves: sharper in taste, bolder in judgment, and more distinctive in how they work.
Thank you to @JoshLu for believing in our bet on open-weight models way back when it wasn’t obvious, and to @_CallMeMacy and @fromevanwang for keeping our feet on the ground and getting us in front of the right people.
Oh, and we’re hiring! If you or someone you know is obsessed with on-device inference, teaching models how people work, or shaping how humans and AI work together, message us.
here's how i shipped 2,500 PRs last month to production
this was originally supposed to be for Cursor Compile in London. i couldn't make it since i was livestreaming for Grok @Bot Galaxy so i'm making it available for free here on X! watch it on 2x speed, i talk slowly
A lot of people have changed their mind about Effect
Why? AI resolved basically every reason people didn't want to adopt it, while making the things it does well even more valuable
@zeddotdev hold on, what about instead of just the 1-level parent, can you do like a n-level call hierarchy. i always find myself annoyingly trying to find the actual ancestor function by function, and this made me think the editor can just show me the entire hierarchy.
2× the number of models trained ≈ 2.02× the parameters in one model.
We fitted a new scaling law with the number of models trained on the x-axis. The exponent came out uncannily close to Chinchilla’s: 0.345 vs. 0.34.
As the population grows, emergent specialization appears and the overall capability of the system improves.
We will provide limited access to the models via API in the upcoming weeks, stay tuned for more.
TLDR:
- Doubling the number of models trained predicts 21.3% lower evaluation loss in our fitted scaling trend. Chinchilla predicts a comparable 21.0% reduction in reducible language-model loss from 2× the parameters and 2.32× the training data. Each model in our population stays the same size.
- At 8 models, selective training used 84% less forward/backward compute, scored 10.7 points higher while saves 87.1% in required memory than training 8 models in isolation.
- At 16 models, the best single model scored 56.0% versus 86.2% for the best-per-task population, a 30.2-point specialization gain
- Emergent skills across benchmarks shaped >80% of specialization, while benchmark identity explained only 1–2%