Publishing a new essay "In Search of a Verifier".
We have seen frontier models make real discoveries in math, code and almost none in robotics, open ended science.
The superficial analysis would be those domains are simply much harder which can be true but I analysed it depends on whether a domain hands you a cheap, exact and automatic verifier.
https://t.co/3sjR3esgYO
well I should abandon my current grad algos if new algos are being invented everyday rendering current lectures old and stale.
Btw the difference is not that significant.
oh yeah btw guys we can do integer multiplication faster than n log n lol
I was definitely very surprised when this one came in lol
https://t.co/o7iBA131nY
Of course, that’s your contention. You’re a first-year machine-learning engineer. You just got finished readin' Attention Is All You Need, probably watched a Karpathy video too. So now you’re convinced everything is just transformers and scaling laws.
That's gonna last until next month when you discover convolution, and then you're gonna be talkin' about inductive biases and locality and how CNNs were actually incredibly compute-efficient for vision.
Then you're gonna read the FlashAttention paper, and suddenly everything's about IO complexity and SRAM and how FLOPs don't matter because the whole goddamn thing is memory-bandwidth bound.
That'll last until somebody shows you an MoE model, and then you're gonna be in here regurgitating DeepSeek, talkin' about expert parallelism and active parameters and how dense models are economically obsolete.
Then six months from now you'll write one shitty Triton kernel, look at an Nsight trace for the first time, and start telling everybody Python isn't actually the bottleneck because the GPU is asynchronous.
And by next year you're gonna be standing right here explaining to me that your 400-billion-parameter model is fast because you turned on CUDA graphs.
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
https://t.co/7N6TPlft1P
great paper and so much to build on top of it.
Agents swarms is the next thing and building the effective transport algorithms, infrastructure is going to be a challenge keeping in mind we have to ensure proper safety, monitoring these systems so all traces can be detected.
Effective coordination is much more trickier than it sounds.
the power of collaboration was so under-rated by me.
Found out yesterday where we had to cook so many items for dinner and together it went so smoothly.
The most profound paper of my PhD so far. We truly did something special to make LMs play/explain chess, needing both architectural & algorithmic innovation (natural-language analogue of the Alphazero algorithm) + applicable to many other domains. Please read on and share!
1/10
@retr0jirachi u are one of the smartest and most hardworking people I have met.
still remember u try to work on ur world model research question without any proper guidance.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
The most important briefings of the day.
Karl Deisseroth, professor of bioengineering and of psychiatry and behavioral sciences, shares the news of his Nobel Prize with his kids.
🚨 MUSE ECOSYSTEM ALERT 🚨
today we are announcing Muse Gadgets!
this is an open-source ESP32 firmware and Linux SDK for anyone to make hardware that works with Muse
we are also releasing our own gadget—Muse Home Link—to enable your muse to work with your smart devices (TV, speakers, etc.)
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next.
We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them.
Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up.
I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay.
I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
After a long wait, releasing the SYNTH paper!
It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency.
Our NeurIPS 2026 main conference paper is out!
We experiment how a simple "verified" source can trip models into agreeing with the answer, while not even trying to verify whether the answer is genuinely correct or not.
We call this Authority Bias!
Paper and website in the comments.
💥 New Paper!
Models are trained not to change their answers just because a user disagrees.
But when the same wrong claim comes from a “verified source,” they give in.
We call this Authority Bias.
Accepted at NeurIPS 2026.
🧵