PhD Student @PKU1898 , focused on RL and Continual Learning for LLM & Agent. Prev Research Intern @Ali_TongyiLab ,@PixVerse_ Now Research Intern @TencentGlobal
Alignment midtraining is one good way of teaching models principles, character, motivations, hopes and dreams. And hopefully, in a way that generalises well. But, this needs to be done while also surviving messy, and sometimes contradictory downstream training.
Some new work on scientifically stress-testing alignment midtraining in this pessimistic regime: https://t.co/ANEQfFI7xa
How we do Data Generation + RL Environment creation are the most interesting sections from recent Open Models releases (DeepSeek & Kimi)
some secret sauce as everyone transforms their data into hill-climbable environments
+ some trends:
1. Model Council + Trace Mining: different intelligence LLMs run on a Task, uncover errors and repair Tasks
2. Strategies for progressively increasing difficulty: more/bigger data, split information across data sources requiring search, multi-modal requirements, cross-domain reasoning
3.Building world structure beyond one-shot generation: Knowledge graphs and event based architectures allow more complicated worlds to emerge naturally
4. Teams have extensive Data Agents per domain with specific verifier logic specific + tips encoded in skills. Ex: computer use vs SWE
every team will benefit from having data research teams that turn every interaction from their agent into a usable artifact for hill climbing
your data is the gold that’ll make your agents better
Everyone moved to GRPO for LLM reasoning because training a critic is a huge nightmare.
Sampling a half-l dozen rollouts per prompt to estimate advantages works, but it absolutely eats through compute.
This new paper pinpoint exactly why value functions keep blowing up in PPO.
It usually comes down to ratio clipping acting weird on low-prob tokens, bootstrapped error accumulation, and GAE breaking on short vs. long responses.
Their fix is a recipe called BPCO (Best-Practice Critic Optimization).
The coolest part is how they handle the critic.
Since you throw the value network away after training anyway, they feed it full reference answers and rubrics that the policy model never gets to see.
Giving the critic "cheat codes" during training stabilizes the value estimates without messing up deployment.
The results look solid. When scaling from 1.5B models up to a 30B MoE, BPCO matches or beats GRPO while sampling just 1 response per prompt.
Definitely worth a look if you're trying to cut down RL compute costs: https://t.co/iLXAXZDKIr
damn sam fires shots at dario and anthropic
"there are some people in the AI field who effectively say we're going to give the world a cure to all disease and we're going to make stuff really cheap in exchange for people giving up their autonomy and impact over the future and power, in the name of safety."
"there will be people that have access to huge amounts of wealth and power and other people just get a pretty good everything. This is a terrible sales pitch."
"the like, dear peasants, we will bequeath upon you these gifts of a cure for cancer and material wealth and great entertainment. And you stop complaining and we'll make all the decisions about the future and just trust us, we'll be better than dictators."
"not good, not good."
In my ~14 months at @OpenAI, one of the most surprising and genuinely wonderful things about it, has been the full leadership support for full-duplex models (gpt-live series), even when it seemed deeply improbable that they would work.
And yet, if they did work, it was obvious how magical they could be. I am quite sure this is not an exception but a general rule for research projects at @OpenAI.
There were so many moments when the problem felt impossibly hard. But one thing kept being true: it was never clear why it shouldn’t work. And almost every time we understood the problem a little better, it became a little easier to solve.
I’m pretty sure there are very few environments in the world where a bet like this could have been made and sustained. So grateful to be part of this wonderful place.
I really do believe we'll see a resurgence in actor-critic style RL by EOY.
Not sure if it'll look more like SAO or something coarser like TEMPO..but so much work this year (BPO, TRACE, OPSD, RLSD, etc) has been circling around the idea of extracting denser training signal from rollouts. A well-trained critic already gives you that natively.
Training and infra lift have been the major blockers to wide adoption of PPO in the past. Serverless training APIs such Tinker have come a long way, though. I'm sure they'll find ways make critic training easier as well.
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
Thoughts About Scaling Law
Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.
The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed.
Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter.
Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it.
This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count.
Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
LE CRITIQUE: Privileged Value Functions for LLM Reinforcement Learning
Value functions for LLMs have been grossly underutilized. I'm excited to finally share my work on Privileged Value Functions, done @MistralAI !
Our paper includes 2 massive improvements to value functions:
1) Privileged Value Functions: Turns out value functions can use hidden context unavailable to the policy for improved token-level RL signal! Condition them on reference solutions, verifier rubrics, or even other group samples and rewards. No more hacky self-distillation necessary! (being a tad hyperbolic)
2) TETHER: Do not trade off group and value baselines, you can just do both! We introduce a new baseline that adaptively trades off between the group mean and value baselines adaptively based on value function reliability.
🌐 https://t.co/i0HfYbVMUX
📄 https://t.co/rPDnf6oOBQ
Code: https://t.co/y54smOwqPz
🧵below
Good blog, makes you think about the empirical observation that cureent RL methods that work for LLMs are *low bias*
- value functions trade off variance with bias, and hasn't shown huge gains yet
- small bias from trainer-inference mismatch often is catastrophic for scaled up RL runs
https://t.co/Ta2DPau1cV
very good blog ackshually to ground your understanding of why and where LLM RL works!
TL;DR:
1. Pretraining induces the strong priors needed for simple policy gradient to work, and RL gives the necessary bits and gradients that we care about
2. Pretraining also reduces the number of bits of info from optimal policy that RL needs to cover
3. Some interesting stuff about SNR and how different techniques help in SNR(SNR roughly scales according to sqrt B in pretraining)
4. Kindof why RL induces a jaggedness in policy and also good explanation of why classical RL even though very sophisticated, was incredibly sample inefficient(game of priors)
Good, easy and short read:)
Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol.
Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented.
A larger context window lets Codex retain more code, tool output, and conversation history before summarizing older material. You need a model that supports it. And GPT-5.6 Sol, for example, has a documented 1,050,000-token window.
Open ~/.codex/config.toml and add or update these settings at the top level, before any [section] headers:
```
model = "gpt-5.6-sol"
model_context_window = 1000000
model_auto_compact_token_limit = 900000
```
The first setting selects the model. The second tells Codex to use a one-million-token context budget. The third starts automatic history compaction around 900,000 tokens, leaving some headroom. Restart Codex client and start a new session after saving.
To try the configuration for a single CLI session without changing your defaults:
```
codex -m gpt-5.6-sol \
-c model_context_window=1000000 \
-c model_auto_compact_token_limit=900000
```
Have fun, but also know that we tuned the default carefully!