We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source.
This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed.
We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition.
Together, we are building an ecosystem of open models that are safe and accessible to all.
A flavour of our research interests on the Dwarkesh Podcast this week, discussed by our very own @oneill_c. Just the start of a longer conversation about long-horizon RL and frontier open-source training here at Base Labs.
New episode with @johnschulman2, @oneill_c and @BerenMillidge.
I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
0:00:00 – Steelmanning the case against RSI
0:18:39 – What’s driving the Chinese labs’ progress
0:28:06 – How will automated AI researchers be trained
0:33:51 – Will long-horizon RL elicit AGI?
0:45:24 – The sim-to-real gap
1:00:33 – How much progress is explained by data?
1:18:03 – Why is RL working so well?
1:24:54 – Move 37 and entropy collapse
1:28:31 – Rapid-fire timelines
Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI.
We believe in a healthy, open frontier model ecosystem. To enable this, we are working on:
- Blue-sky research on continual learning, the science of RL, and how models learn across their full lifecycle, with every experiment and recipe shared openly.
- The BaseHub Data Foundry: the highest-quality open RL environments, training data, and real-world benchmarks, built for anyone to train and benchmark on.
- Post-post training: taking open-source models and making them better, safer, and more aligned through continual post-training, built on our research, and deployed with our frontier safety stack so organizations can use open-source models with confidence.
- Making models cheaper and more performant through our model performance research.
This is a mission-driven research effort, not a commercial product. We believe the health of the open-source AI ecosystem matters and that the best way to advance it is to do science in the open.
We’re hiring engineers, researchers, and research fellows to advance this mission.
https://t.co/Go8YsfaP2X