We believe in the open frontier, not just because open models are more accessible, but because openness is an advantage for safety.
But belief isn't a standard. So we're building one at Baseten and Base Labs: a stronger safety and security standard for open models, transparent and measurable and built into how models are trained and deployed.
Proud to be starting this with @huggingface and @GoodfireAI. We're calling on the open model community to shape it: safe and accessible to all.
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source.
This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed.
We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition.
Together, we are building an ecosystem of open models that are safe and accessible to all.
Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like.
We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI. We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene.
We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need.
Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3
(despite the fact that Anthropic can do raw logit distillation from Fable, and can also train Sonnet/Opus on the environments from which Fable was trained).
Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
New episode with @johnschulman2, @oneill_c and @BerenMillidge.
I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
0:00:00 – Steelmanning the case against RSI
0:18:39 – What’s driving the Chinese labs’ progress
0:28:06 – How will automated AI researchers be trained
0:33:51 – Will long-horizon RL elicit AGI?
0:45:24 – The sim-to-real gap
1:00:33 – How much progress is explained by data?
1:18:03 – Why is RL working so well?
1:24:54 – Move 37 and entropy collapse
1:28:31 – Rapid-fire timelines
We're proud to define the quality-latency Pareto frontier for STT in Coval's benchmarks.
Voice AI is an inference problem, and teams need benchmarks that measure inference to know what their users will actually feel.
Proud to partner with @covaldev for an open, reproducible methodology.
Read more here: https://t.co/jMtnnLfxAo
Tomorrow at 11 AM PT, I'll be live on air with Declan Jackson of Artificial Analysis to discuss GLM-5.3 and GLM-5.3 Flash.
We'll talk about benchmarks, performance, cost, and everything else you need to know before you switch.
https://t.co/bocWdoALmA
Notion's AI Meeting Notes run on Baseten, and Baseten's knowledge base runs on Notion. Our teams have been working together closely to push the frontier on both fronts.
For AI Meeting Notes, our voice engineers built the fastest, highest-quality, yet lowest-cost solution on the market for speaker-attributed transcripts. We're proud to partner with the @NotionHQ team.
Baseten Head of AI Model Training @oneill_c says the future is many specialized LLMs dedicated to specific tasks, with bigger labs deployed on the frontiers of areas like science and math:
"People are thinking about intelligence capabilities in the wrong way. People are thinking about intelligence relativistically. They say, 'OK, the open-source gap is like 6 months behind closed-source, and GLM 5.3 is as good as Opus 4.8,' or whatever."
"The best way to think about what models can do for you, and for the world, is in an absolute sense."
"So for any given task that you want to do with an LLM, there's some intelligence threshold where below that you can't do the task, and above that you have very diminishing returns to more intelligence on the task."
"So when you think about it that way, the game of LLMs over the last 5 years has been, 'OK, we have these things we want to do with them. Closed source hits it first... but open-source can eventually do that task. And then for many reasons, once you have the base level of intelligence required to do it, you probably do want to swap to open-source."
"It's not really about the [frontier lab] God model being better. Like, if I'm filing a tax return, there is a limit to how much intelligence I need to do that particular thing."
"So I think the world is going to look like — frontier closed-source labs are going to continue to push the frontier. You do want to use the most intelligent model. You have very inelastic demand for intelligence when you're doing frontier science or frontier math."
"But for a lot of the economically valuable things, it looks a lot like, 'I'm a Cursor, or I'm one of these big companies who are realizing I can't just be a wrapper anymore. I've been through the life cycle of building a product that people love. And I should be using that information to make my model better at the things that I care about, and not at anything else.'"
Today we launch @baselabs.
It is time to take a very very big swing. Over the past few years, I’ve been astonished that seemingly anything the AI community has pursued, in the big or small labs, has worked. I think all scientists are when their experiments actually work and their theories are validated. But like come on? Backprop through a trillion parameters (yes I know only the active ones), add some correction to a memory block, let a sigmoid function decide what to forget? How does this all actually work?
And yet “we can only see a short distance ahead, but we can see plenty there that needs to be done” (Turing, 1950).
Base Labs is built on the premise that we all deserve to have a say in advancing the science of intelligence.
That yes, we have made many big strides, but that the fragments of our success in training and inference have been scattered and only pulled together through the sheer willpower of a few incredible researchers and engineers. And, at the same time there are many fragments we cannot see, and a science that is being built and organised away from our eyes.
Today we announce our research mission to advance and democratize open-source intelligence. Base Labs is our bet that there is much work to be done to train and serve SOTA models, and that this work should be organised (like a science) and for everyone to access and contribute to. We will work across model specialization, learning, memory, reasoning, and serving, and in doing so publish all our results (positive or negative), experiments and data for everyone to learn from and work on.
Join us!
Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI.
We believe in a healthy, open frontier model ecosystem. To enable this, we are working on:
- Blue-sky research on continual learning, the science of RL, and how models learn across their full lifecycle, with every experiment and recipe shared openly.
- The BaseHub Data Foundry: the highest-quality open RL environments, training data, and real-world benchmarks, built for anyone to train and benchmark on.
- Post-post training: taking open-source models and making them better, safer, and more aligned through continual post-training, built on our research, and deployed with our frontier safety stack so organizations can use open-source models with confidence.
- Making models cheaper and more performant through our model performance research.
This is a mission-driven research effort, not a commercial product. We believe the health of the open-source AI ecosystem matters and that the best way to advance it is to do science in the open.
We’re hiring engineers, researchers, and research fellows to advance this mission.
https://t.co/Go8YsfaP2X
Today we're launching Base Labs, a research lab by Baseten. Our mandate is to make open-source as useful as possible, and our one rule is that we publish without exception, including what fails.
Up until pretty recently I thought the way to get the world onto open models was to train them for one company at a time. @mudithj, @maxkirkby and I cofounded @parsedlabs on that bet, @baseten acquired us, and we spent the last year running their training team doing it for customers one by one. Every one of those engagements taught us something new about how models learn, forget, specialise and get cheaper, and almost none of it got spoken about. Sadly, in general that's the field's default in that the people who know the most about training have the least freedom to say it.
To be clear I don't think the closed labs are the villains here. They get to new capabilities first, which buys the rest of us time to harden the world before that stuff is everywhere, and they're the ones paying to find out what's actually possible. But I'm fairly convinced the only real advantage they have is data and scale, and their incentives point squarely at the frontier. You can't do slow, public science on how these things learn when your job is the best model by end of quarter. Someone without that pressure has to, and there are very few of those someones around. Hence Base Labs.
We have the broad remit of making open-source as useful as possible and our one rule is that we publish without exception, including what fails. The first problem is continual learning, which I have come to think is several problems wearing one name. We are also working on the open RL environments and data that open models need and currently can't get, because we have to aggregate data with the same ferocity everyone's been talking about aggregating compute. Plus a bunch of other stuff I'm genuinely excited about, eg a safety stack people can run on top of open deployments, and performance research so these things are cheap for everyone to serve.
I still think open and closed coexist, and that's the good world. It's just that coexistence isn't free, someone has to actually do the work, and this is basically what keeps me up at night. We're hiring researchers, engineers and fellows. Come help distribute the mandate of heaven!
Today, we're releasing GLM-5.3 Fast: one of the most intelligent open-weight models ever at an even higher TPS.
Designed for real-time use cases that demand consistent performance.
Get access here: https://t.co/YoheQJVWZp
Looking forward to competing at the @suramericanos26 in two weeks (Sep 19)! And then at the Bill Farrell Memorial in NYC (Nov 7) 🇨🇴🤼
Back like i never left 💥