🚀New Paper!
https://t.co/pyiwfwTkuL
While fact verification is essential to ensure the reliability of LLMs, detailed analysis of fact verifiers remains understudied.
We present several findings based on our revised dataset, along with practical guidance to improve the models.
I will be at @COLM_conf this week!
I'm interested in agentic envs for RL, and scaling synthetic data to improve LMs. I'm also continuing research on Deep Research agents in enterprise scenarios as part of my internship @Microsoft .
Please reach out to chat or just grab coffee!
Many RL algorithms are used to train LLMs, but if we trace their evolution there is a clean progression...
VPG → REINFORCE → PPO (Actor-Critic) → GRPO → GRPO variants.
VPG. We can begin with the vanilla policy gradient (VPG), which provides the foundation for nearly all policy gradient algorithms. Intuitively, VPG—and policy gradient algorithms in general—aims to increase the probability of actions that lead to good outcomes (i.e., high rewards) and decrease the probability of actions that lead to bad outcomes (i.e., low rewards).
The VPG can be derived by taking the gradient of a standard RL objective. It yields the familiar policy gradient expression we see in most all RL algorithms: the gradient of the log-probability of an action is multiplied by some learning signal, such as a return or advantage.
REINFORCE. The VPG is written as an expectation over trajectories, which we cannot compute in practice. To solve this, REINFORCE instead samples a finite set of rollouts, computes the policy gradient expression for each, and averages across them—giving us a sampled / Monte Carlo estimate of this expectation. This allows us to actually compute this expression in practice, but such a Monte Carlo estimate tends to have high variance. To reduce variance of the policy gradient estimate, we can subtract a baseline from the return.
Actor-Critic. Although simple and practical, this Monte Carlo approach can still produce high-variance policy gradient estimates. To further reduce variance, we can use a more informative baseline: the value function.
This motivates the actor-critic framework, where we simultaneously train the policy—or actor—and a value model—or critic—during RL. The critic estimates the expected return from each state, and we can compute the advantage by comparing the return after taking an action to this expectation. Intuitively, this tells us how much better or worse an action performed relative to what was expected.
PPO is one implementation / example of such an actor-critic framework. It combines powerful advantage estimation with conservative policy updates. PPO samples rollouts and can perform multiple policy updates over this same data. As the current policy changes, an importance ratio accounts for differences between the current and old policies, where the old policy is the policy that was used to sample the rollouts. PPO uses a clipped surrogate objective that limits the benefit of moving this ratio too far from 1, preventing excessively large or destructive policy updates.
GRPO. The downside of actor-critic algorithms is their overhead: training a critic introduces non-negligible additional memory, computation, and complexity. Aiming to solve this problem and create a more straightforward RL algorithm, GRPO removes the critic entirely. Instead, we sample several completions for the same prompt and estimate the advantage of each completion based on its reward relative to the rest of the group. We retain much of the PPO-style clipping and objective without needing to train a separate value model.
GRPO variants. The story does not stop with vanilla GRPO. Many more recent methods (e.g., DAPO, Dr. GRPO, GSPO, TIS, CISPO, and more) build upon the same general policy-gradient foundation. Rather than completely replacing this underlying framework, these methods modify details like loss aggregation, reward normalization, sampling, clipping, importance ratios, and the handling of off-policy data.
This is one of my biggest takeaways from studying RL for LLMs: there are many algorithms, but the conceptual foundation underneath them is surprisingly coherent. If we understand core concepts like the VPG, advantage estimation, importance sampling, and constrained policy updates, most modern RL + LLM research is much easier to grasp!
John Schulman (co-founder of Thinking Machines and OpenAI) wrote that goal-driven research beats idea-driven research.
I realized it's also really good life advice.
Don't chase what other people are doing. Figure out what life you want, and then figure out how to get there.
Sharing my first of hopefully many research blog posts! https://t.co/qPvLWmsghr
This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs.
I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.
New mini-paper: "A note on goal-based hierarchical RL". It combines the cool agent-centric general value function (ACGVF) construction of @geraudnt with work I did 25 years ago (!) on hierarchical HMMs. Caveat: no experiments yet... https://t.co/gcLHHQks8S
RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (https://t.co/uniFmpbDOn). Now, we propose TailSFT (https://t.co/75yejb8voI), a lightweight + principled way to directly improve coverage and get better post-RL perf.
How does one RL post-train a 397B model for long-horizon knowledge work? 👩💼
We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research.
Full blog: https://t.co/GWgR2DB7iy
Source code: https://t.co/96h8ADfvBp
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.
Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.
Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Mental model on why RL/GRPO is a bet on Search & Synthetic Data
for a given task, the “correct trajectory” lives somewhere in the model’s output space
“good models” are then good because they can sift through the infinite possible trajectories to reliably output the correct trajectory for a set of tasks
or one of its isomorphisms (as long as the verifier captures that isomorphism)
every rollout in GRPO is literally a piece of synthetic data produced by the model
so the model is able to eventually find the right trajectory itself some of the time instead of always requiring a human labeler to write the exact steps
this has another interesting side effect:
sometimes models can discover more efficient solutions simply by scaling the search policy conditioned-> ie. do more rollouts and reward more efficient solutions
human labeling CAN be a local minimum -> see AlphaZero
and every RL step is adjusting the search policy conditioned on the set of Tasks that are available during training
so
- if you make good Tasks
- have a good harness that can allow the model to solve the task some of the time
- then you can use RL to refine the search process and discover the correct solution more reliably (eventually higher pass rate)
What's left for humans in a world where machine intelligence has so many advantages?
I recently got a Tesla, and using full self-driving has been a wake up call to just how many advantages AI has over humans. The few times I disengaged it because I thought it was going into the wrong lane, it turned out that the car was right and I was wrong. I realized that there is no hope of me driving better than a neural net that knows every road, sees in every direction at once, and never gets tired or distracted.
Given that AI has certain inherent advantages over human intelligence, what kind of moats will remain for us as humans? It's a big question.
One short-term answer is that the world we live in was created for humans, and in some domains, AI has not closed the gap yet. For instance, AI still struggles to use internet user interfaces. While any computer-literate human can navigate a web page with ease, AI is still not great at making accurate clicks and drags because image embeddings are not optimized for such precision. If the internet were designed to be fed into language models instead of rendered as visual interfaces for humans, AI would obviously far exceed humans. But for now, language models still need to be retrofitted to our legacy infrastructure.
Robotics is another area where we humans have a home-field advantage. Most tasks in the physical world are designed around fingers and opposable thumbs, which have been pretty hard to build into robots so far. While it is clear that machines can outperform humans in environments optimized for automation, like large-scale manufacturing lines, for now, most of the world is still built for humans.
However, these capability gaps are only temporary. There will surely be a day when machines click faster than us and have superhuman general dexterity. What are the real moats that humans will have?
Anything involving private knowledge that language models do not have access to feels like a solid moat to me. Romantic matchmaking and high-end real estate are two examples where inventory is often not advertised publicly and matches are made through being in the right circles. Venture capital is another example—although some research and decision making can be automated with AI, much of success hinges on understanding trends ahead of time and connecting the right people, both of which require private knowledge. While machines can and probably will have increasing access to some types of private knowledge, I think there will still be some types of private knowledge that only humans know. I do not see a path for AI to win when critical knowledge is closely guarded in human circles.
A second area where humans seem to have a real moat is in entertainment and the arts, which are inherently valued for their human aspects regardless of how well machines can do them. Watching Usain Bolt sprint one-hundred meters is beautiful as an expression of the peak of human ability, even though cars can drive much faster. Watching chess at the amateur or intermediate level is more relatable and satisfying than watching two superhuman AIs play each other. The value of art comes from the creation process, which is why replicas are not as valuable as originals. These types of work feel like they will continue to have markets even as we advance towards superintelligence.
More broadly, human presence is a feature that will be, by definition, challenging for AI to automate. For example, a teacher remembering your name or a parent supporting you is valuable even though AI can easily remember your name and probably give better life advice. Someone spending part of a finite life on you counts because their time runs out. As a personal anecdote, I remember the first time I worked with someone who I considered an amazing AI researcher. His advice was solid but what was more important was that I believed I could do great work with him as a collaborator and I raised my own standards. Over the past decades, the development of technology has divided us in some ways, but hopefully AI brings us closer to a world where human presence is reemphasized.
Intelligence has been the defining feature of humans and it will be a big change for AI to automate that over the coming decades. In the near term, certain types of intelligence will become very cheap and automate away old jobs, but the moats I described above will not be the only places where humans can hold value. In the same way that computers took away the jobs of secretaries and manual accountants but created far more jobs via the IT industry, I believe there will be much more demand for services created by productive AI-augmented humans, perhaps for services we cannot yet imagine in today’s society. Just seeing how this story plays out will be an adventure in its own right.
New paper! LLMs Corrupt Your Documents When You Delegate
LLMs are enabling a new way of working: delegated work, where users supervise an LLM as it edits documents on their behalf.
Delegation requires trust: does the LLM complete tasks without introducing errors?
We simulate delegation across 52 professional domains and find that LLMs Corrupt Your Documents When You Delegate. 🧵1/N
we need agent evals that are really consistent with real world usages. otherwise people are optimizing foundation models for the wrong direction. the problem of targeting is even bigger than benchmaxxing.
New OpenAI post: Can midtraining on docs about aligned AI bake in alignment priors for agents? We report an experiment where those priors are quickly washed away by RL and fail to generalize to agentic settings. But that cuts both ways: priors that AIs are misaligned fade too!
can synthetic training beat RAG in data-constrained domains?
we suggest a simple recipe for better synthetic training:
- Synth Mixed Training: train on both synth QAs and synth docs
- Focal Rewriting: rewrite docs with targeted topic prompts
results:
- beats RAG by +2.6% on QuaLITY
- improves to +4.4% with Focal Rewriting
- reaches +6.7% when combined with RAG
Paper: https://t.co/0rnfqUH5nE
One gem from Composer paper is that RL improved both pass@k & pass@1. Suggests RL does not just reweigh existing capabilities but also teaches new ones? 💎