LLM Post-Training/ Recursive Self-Improvement for Reasoning
Co-founder/ MTS @ReasonCoreAI
PhD @UCLA, Bachelors @iitbombay
Led a team of ~500 PhD researchers
One minor reason I believe data will be enduring is because, while we will and are hillclimbing on benchmarks that make models more tasteful, taste is a moving target. A seemingly unserious but important example of this is that agents are still very bad at planning great experiences (eg trips, events, restaurants) even though much public information exists about these things. by the time taste becomes publicly known and consensus, taste has moved on.
The model that produces the best attempt doesn't have to be the model that recognizes it. In a new Terminal-Bench 4 experiment, a cheaper DeepSeek model improved results by selecting among Claude Opus attempts. That doesn't make it a better coding agent, and generating several attempts still costs money. It suggests a useful training target: recognizing when the work is actually finished, rather than learning to sound finished. https://t.co/1VVNh8GQ4L
Evaluating AI data providers: Beyond benchmark-maxxing
Benchmark-maxxing gets an AI data vendor the first sale. Labs are sophisticated buyers. Every serious lab keeps held-out internal evals the vendor never sees. It runs an ablation on each data purchase and compares the gain to the price. Data that only lifts the vendor’s own benchmark looks like noise in that ablation. The vendor gets a first purchase order and never a second.
Labs pay for 5 things, roughly in order of how long the spend lasts.
1. Fixes for failures they see in production. A lab knows where its model breaks for paying users from usage logs and enterprise escalations. A coding agent loses track after 40 steps. A finance workflow makes up a number in the third tab of a model. A vendor who can take a cluster of failures and return data that fixes it will get the next contract too.
2. Expertise the lab can’t generate cheaply. Synthetic data and model-generated traces cover a lot now. They can’t reproduce how a tax attorney or a staff engineer reviewing a 2,000-line diff makes judgment calls. Labs pay for access to these people, and more and more they want the expert’s reasoning and rubric along with the answer.
3. Environments and graders for RL. As post-training moves toward reinforcement learning, labs buy environments: a sandboxed task, a way to check whether it succeeded and a reward signal that’s hard to game. A good environment keeps producing training data long after the vendor delivers it. My read is that this is where the biggest checks are going now.
4. Measurement. Labs buy private, uncontaminated eval sets because public benchmarks have leaked into pretraining data. So the vendors gaming public benchmarks are part of why labs need private ones.
5. Speed and exclusivity. An in-house expert network takes a year to build, and buying one takes a month. A lab will sometimes pay extra for exclusivity just to keep a competitor from training on the same data.
If you’re evaluating a data company, 4 questions separate real lift from benchmark-maxxing:
1. Does the gain show up on the lab’s internal evals, or only on evals the vendor built?
2. Does it carry over to related tasks the data didn’t target?
3. What’s net revenue retention by lab? If each lab buys once, that’s benchmark-maxxing. Growing contracts mean the data works.
4. Who defines the problem? If labs bring their failures to the vendor, that’s pull. If the vendor shows up with a benchmark and a pitch, that’s push.
The AI data companies that last will be the ones labs bring their problems to and that fix production model failures or help them hill climb internal evals, not just external benchmarks.
If you'd like to learn about Harbor, Terminal-Bench, and Terminal-Bench-Science without reading through a wall of text, check out these lightning talks by @alexgshaw, @ryan_marten, and me from our recent @terminalbench meetup. Now live on YouTube📽️
I’ll be in SF all day on Thursday & Friday this week. If you’re in SF for COLM or SF Tech Week and would like to chat and get Terminal-Bench-Science merch, get in touch via DM!
Thanks to @LaudeInstitute & @harborframework for organizing the merch.
Always loved the work done by @ArtificialAnlys - excited to hear that Terminal-Bench-Science will be included in the next version of the Artificial Analysis Intelligence Index 🚀
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇
With Muse and Dots recently launched by Meta and OpenAI we’re happy to share our new survey on Proactive AI!
This is a great collaboration with co-authors from UIUC, Google, Berkeley, and CMU. We discuss how AI systems can move beyond reactive assistance toward proactively anticipating needs, making decisions, and taking actions—while carefully considering the risks that come with greater autonomy.
Excited to see where proactive AI goes next! 🚀
📃 Paper: https://t.co/JC8VThc3GA
a core prediction i made to friends and founders who i work with almost 10 months back was that all data companies will end up doing research and it will be a core part of their business. great to now see the best ones actually share how they are doing this (they are best in the business for a reason instead of the barrage of slop data selling companies)
the best ones will move beyond just post-training Qwen 4Bs to show hillclimbable-ity. they will work on research problems like designing calibrating on true open ended long horizon tasks (cc universes), simulators, automated verifier expansion and fuzzing, adversarial pipelines for mining core failure patterns, env design for higher order capabilities like self-verification in task domains, creative task specific reward functions (reconstruction rewards, etc.) for better signal and more.
in particular, verifier design specifically has so much research headroom and is very hill climbable in my opinion if a domain expert can encode objectives (e.g. autoformalization for software), and in general more ideas on posttraining models themselves to improve on synthetic env generation (the best researchers are already working on this and the credit is to them for this line of thinking)
people dismiss data companies very quickly, but little do they know that the great ones tackle good and useful research problems
We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports.
Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping.
All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September.
Read the full report: https://t.co/mGeVIBgdou
Come checkout our poster SUPERNOVA: Eliciting general reasoning in LLMs with Reinforcement Learning on Natural Instructions!!!
TL;DR: instruction datasets are an underused RLVR source, but utilizing them requires principled empirical curation.
📍 Imperial Ballroom B #COLM2026
Check out our latest paper on looped transformers! We found looping enables implicit reasoning, and a new trick we introduced that forces the model to decode the tokens explicitly between looped modules enables the model to learn multi-hop reasoning in a generalizable way!
From predicting the next token to anticipating the next moment.
Coming from NLP, I’m excited to share my first humanoid robotics paper, Reflex! 🤖
I believe the next frontier of AI is not just understanding the world, but perceiving, reacting, and acting in it in real time.
What I read at #COLM2026: the 5 papers I liked most, ranked by community likes:
1. IdeaScientist: agents trained with RL to generate grounded research ideas. @Jiarui_Liu_
https://t.co/CM0N8qBqgI
2. LLM-as-a-Verifier: a weaker model checking a stronger one’s work. @Azaliamirh
https://t.co/1I6P5e63kI
3. CORAL: autonomous multi-agent evolution for open-ended discovery. @ao_qu18465
https://t.co/yitgWGuYVA
4. Actor-Curator poster. @lightetal
https://t.co/yuYkafFV1Y
5. GitSwarm: inference that builds on itself across long tasks. @veds_12
https://t.co/hhIJh0tA2t
My overall impression is that verification is now the bottleneck for both agents and reasoning. Which COLM paper did I miss?
Excited to share IdeaScientist, our research from my internship at Meta! 🚀
Can AI agents learn to generate novel, grounded scientific ideas?
We introduce IdeaScientist, a framework that trains specialized agents to identify research gaps, discover useful connections across scientific domains, and turn them into concrete research proposals.
🔍 Key findings:
• +14% overall performance over open autoresearch baselines
• 25% improvement in novelty, showing that explicit training can improve scientific ideation
• 75–92% human preference win rates against six open baselines in blind evaluations
• Cross-domain retrieval substantially improves the transfer of ideas between research fields
We also introduce Svalbard Idea Vault, a collection of 2.77M decomposed research ideas for training and evaluating scientific ideation.
📄 Paper: https://t.co/OCFIyCKe4k
Native VLA/WAMs have very weak language-grounding and reasoning capabilities. With ARC, we uplift existing VLA/WAMs to gain ~30pp performance with no new robot action data. The capability is already in these models; we just haven’t fully unlocked it yet.
https://t.co/u3WOH9ErzK
An RL environment is a task and a check.
The task takes a day to write. The check takes weeks, because the agent will find every path to a pass that skips the actual work.
Most of building environments is closing those paths.
Parallel attempts repeat work. Long reasoning chains lose detail through compaction. How can long-horizon inference COMPOUND? 📈
GitSwarm: Decentralized Compounding Inference.
A persistent memory swarm that preserves and builds on work in a shared Git repository. 🧵
[1/n]