Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next.
We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them.
Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up.
I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay.
I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
The Pareto frontier is misleading ❌ for general users. Price it by subscription instead of API. With cache hits, Opus 5.5 with Claude Code 20x has already killed every other model.
I can't find the exact video, but @jeremyphoward put this perfectly and simply in part 2 of the fastai course: when you come across something you don't understand, look it up, read about it, ask questions, understand it, and move forward. Repeat.
There's no need to think you're stupid. There's no need to beat yourself up about being ignorant. You don't know something? --> look it up --> fill the gap --> now you know. Repeat.
That loop is at least multiplicatively if not exponentially more powerful with LLMs now.
Today we establish that AI is as good as expert human tutors for immediate GRE learning gains (p=.015, n=2,383 students, study conducted July-Sep 2026).
One hour with an expert GRE tutor: $75.
One equivalent hour with an AI tutor: 7 cents .
918x cheaper, equivalent learning (p=.044, n = 140).
In 5 of 7 academic topics, the top AI tutor beat the expert human tutors on average.
How we measured it: pre-test -> 1 hour with (AI or human or no) tutor -> post-test.
There are no LLM judges. Learning gains are post-pre. Everything is real students.
Why we did this:
> 1 trillion USD is being spent per year on machines making machines better. StudentBench shows us how to shift these resources so machines also make humans better.
Why I did this:
I grew up in rural Kentucky. My dad, grandad, and great grandad were all mailmen. Getting into a good college changed my life. Now we have done a rigorous study to help AI labs evaluate LLMs and help show how bright AI can be for students of all incomes.
I believe in good science that leads to the democratization of opportunity.
The research paper behind StudentBench: https://t.co/JfkDDllP4i
Data, code, details: in this thread
If you're wondering what I'm selling, sorry to disappoint: we released all study data for free on Hugging Face, and made https://t.co/5G9L8yAxrt free to use. Enjoy.
lots more in the video / thread!!
@ben_golub Always before, in mathematics, solving a problem was nearly entirely aligned with finding useful reusable teachable ideas that pushed the field forwards.
Now that's not always true any more.
I think that's a really important and interesting insight! :)
@ben_golub The letter really resonated with me. Your point "if your famous problems aren't ones that you really want to the answers to, something's wrong" is also the point made in the letter, at least to some extent.
At least n=2 but it’s not a lot! (synbio for two decades: dna synthesizers, sequencers, cell engineering, viral design and have been scaling transformers at GDM since 2018.) I agree with all of your points - thanks for making them. The biorisk pandemic stuff is straight out of the crackpipe from hucksters who clearly never worked in a lab. No appreciation of the timescales, biophysical complexity, supply chains, or economics.
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands.
And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
This is the type of post I might regret later, but I really had an urge to rant about the current discourse around pacing the frontier. https://t.co/A0Vuk4DXEF
It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs.
So today we're launching the Open Alignment Initiative, led by @Thom_Wolf@huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to.
Let's make AI safer by making it more transparent!
Today we're launching Base Labs, a research lab by Baseten. Our mandate is to make open-source as useful as possible, and our one rule is that we publish without exception, including what fails.
Up until pretty recently I thought the way to get the world onto open models was to train them for one company at a time. @mudithj, @maxkirkby and I cofounded @parsedlabs on that bet, @baseten acquired us, and we spent the last year running their training team doing it for customers one by one. Every one of those engagements taught us something new about how models learn, forget, specialise and get cheaper, and almost none of it got spoken about. Sadly, in general that's the field's default in that the people who know the most about training have the least freedom to say it.
To be clear I don't think the closed labs are the villains here. They get to new capabilities first, which buys the rest of us time to harden the world before that stuff is everywhere, and they're the ones paying to find out what's actually possible. But I'm fairly convinced the only real advantage they have is data and scale, and their incentives point squarely at the frontier. You can't do slow, public science on how these things learn when your job is the best model by end of quarter. Someone without that pressure has to, and there are very few of those someones around. Hence Base Labs.
We have the broad remit of making open-source as useful as possible and our one rule is that we publish without exception, including what fails. The first problem is continual learning, which I have come to think is several problems wearing one name. We are also working on the open RL environments and data that open models need and currently can't get, because we have to aggregate data with the same ferocity everyone's been talking about aggregating compute. Plus a bunch of other stuff I'm genuinely excited about, eg a safety stack people can run on top of open deployments, and performance research so these things are cheap for everyone to serve.
I still think open and closed coexist, and that's the good world. It's just that coexistence isn't free, someone has to actually do the work, and this is basically what keeps me up at night. We're hiring researchers, engineers and fellows. Come help distribute the mandate of heaven!
@Yuchenj_UW I'd love to try it, thanks :)
I tried googling for info about models in Databricks and also reading some docs, but I found it totally inscrutible.
I'm in San Francisco this week. So is the Deputy Prime Minister, but we’re here for very different reasons.
He's meeting with Anthropic and OpenAI, working on Australia’s data centre pipeline.
I'm here as the co-founder of @_Firmable and a Partner at @GlitchCapital, meeting with Australian founders who moved here this year. The CGT changes came up in almost every conversation as part of the reason they left Australia to build their business in the US.
What a mess.
The AFR published two opinion pieces todaythat demonstrate this shambles.
In the first, Treasurer Jim Chalmers set out how Australia can benefit from AI if we get it right. And that’s a big if. The headline number he was spruiking was $150 billion of data centre investment by 2030.
Jim, data centres are the bottom of the stack. Concrete, power and imported GPUs. Enormous capex, very few jobs, and negligible compounding growth for our country. If that's only our AI strategy, we’re in trouble, as we're importing expensive chips, hosting someone else's product, and then buying the intelligence built on top of it at full retail.
The bigger opportunity is in using and complementing the frontier models. Products built on top of frontier models that solve real business problems. That's where jobs are created and where the productivity Chalmers is chasing actually comes from. I’m seeing it at what we’re building at Firmable, the companies we invest in at Glitch and the founders I’m meeting in the Bay Area.
In the second op-ed, @Shaun_Cartoon laid out what Chalmers is actually doing to the people who build that layer. Effective rates for founder and employee equity are doubling to 47%, which is why founders are leaving.
Australia is courting the frontier labs on one hand, and taxing the hell out of the Australian founders and builders.
You can't be a value-add exporter of AI when the value-adders are in San Francisco.
If Australia is serious about backing innovation, restore the 50% CGT discount for all businesses. This alone will generate far greater productivity gains for our country than data centres will.