Nice to see this idea getting attention!
This is essentially what we called "depth-wise batching": overlapping requests at different recurrent depths (much higher utilization with adaptive compute).
Recursive Models: https://t.co/abCjYgcfGx
MoR: https://t.co/q2BGHi6f77
๐ข๐ป-๐ฝ๐ผ๐น๐ถ๐ฐ๐ ๐ฑ๐ถ๐๐๐ถ๐น๐น๐ฎ๐๐ถ๐ผ๐ป ๐ถ๐๐ป'๐ ๐ฎ ๐ณ๐ฟ๐ฒ๐ฒ-๐น๐๐ป๐ฐ๐ต
On-policy distillation has become a default post-training tool in many open-source frontier model training recipes. Recent releases lean on it heavily: DeepSeek v4, MiMO, and Nemotron-Cascade-2 use MOPD, and GLM 5.x uses on-policy cross-stage self-distillation. It provides RL's on-policy nature reducing exposure bias, while providing token level supervision like SFT.
But OPD and OPSD have their own failure modes. In this post I discuss a few of them:
1. ๐๐ฎ๐ฟ๐น๐ ๐บ๐ถ๐๐๐ฎ๐ธ๐ฒ๐ ๐ฎ๐ฟ๐ฒ ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฎ๐น๐น๐ ๐๐ป๐ฐ๐ผ๐ฟ๐ฟ๐ฒ๐ฐ๐๐ฎ๐ฏ๐น๐ฒ. When the student samples a rollout and takes an early wrong turn, the per-token KL computed along that frozen rollout cannot pull it back onto a correct path. TRD proves that this failure is built into the objective rather than being a matter of noisy gradients. Even with a perfect teacher, the gradient obtained from token-level KL on the student's own rollout agrees with the ideal corrective gradient at exactly one point, the token where the student first diverged, and disagrees everywhere after it. Every later supervision target is therefore anchored to a context that the student should never have entered. Because reweighting or clipping only rescales the magnitude of each token's gradient, and here the terms point in the wrong direction, no per-token adjustment can recover the correct update. TRD's proposed fix is to distill along a teacher-refined trajectory rather than the raw student rollout, which restores a target the student can actually follow.
2. ๐ ๐๐๐ฟ๐ผ๐ป๐ด๐ฒ๐ฟ ๐๐ฒ๐ฎ๐ฐ๐ต๐ฒ๐ฟ ๐ฐ๐ฎ๐ป ๐ฏ๐ฒ ๐ฎ ๐๐ผ๐ฟ๐๐ฒ ๐๐ฒ๐ฎ๐ฐ๐ต๐ฒ๐ฟ. On-policy distillation can only teach the student at states the student itself visits, and the usable signal at each of those states lives in the overlap between the student's and teacher's next-token distributions. Rethinking OPD shows that a bigger, higher-scoring teacher can fail to move a student while a weaker one succeeds, because if the teacher's token distribution places its mass on tokens the student rarely produces, the overlap is small and almost nothing transfers, no matter how capable the teacher is in absolute terms. What actually predicts success is early top-k thinking-pattern overlap. In runs that work, the shared top-k tokens carry 97 to 99% of the probability mass and the overlap ratio climbs steadily during training, whereas a run that starts with low overlap never recovers it. A teacher trained on the same recipe as the student also converges toward the student's own distribution, so its higher benchmark score does not correspond to any new knowledge it can transfer. The practical rule is to pick teachers by distributional closeness to the student, not by leaderboard rank.
3. ๐ฃ๐ฟ๐ถ๐๐ถ๐น๐ฒ๐ด๐ฒ๐ฑ-๐ถ๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป-๐ฐ๐ผ๐ป๐ฑ๐ถ๐๐ถ๐ผ๐ป๐ฒ๐ฑ ๐ข๐ฃ๐ฆ๐ ๐ฐ๐ฎ๐ป ๐ณ๐ฎ๐ถ๐น ๐๐ผ ๐๐ฟ๐ฎ๐ป๐๐ณ๐ฒ๐ฟ. In OPSD you distill a teacher that was conditioned on privileged information, such as the gold answer, into a student that will never have it. The Many Faces of OPD shows what goes wrong when that information is instance-specific. The student cannot recover the teacher's per-instance reasoning, since it never sees the answer, so it instead learns a single answer-free policy that effectively averages the teacher's behavior across all problems, and that averaged policy is too generic to solve any particular one. The signature is initial gains followed by collapse: rollouts grow long, fill with hedging tokens, and accuracy craters toward zero. The approach works only when the privileged information is a shared rule that applies across all instances, such as a system prompt or an alignment preference, and not when it is a per-problem answer.
4. ๐ง๐ต๐ถ๐ป๐ธ๐ถ๐ป๐ด ๐ฐ๐ผ๐น๐น๐ฎ๐ฝ๐๐ฒ: ๐ฑ๐ฒ๐ป๐๐ฒ ๐๐๐ฝ๐ฒ๐ฟ๐๐ถ๐๐ถ๐ผ๐ป ๐๐๐ฝ๐ฝ๐ฟ๐ฒ๐๐๐ฒ๐ ๐๐ต๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น'๐ ๐ผ๐๐ป ๐ฑ๐ฒ๐น๐ถ๐ฏ๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป. A teacher conditioned on the answer has no reason to hesitate, backtrack, or explore, so its per-token targets quietly push down the student's deliberation tokens. Diagnosing and Mitigating Thinking Collapse names this phenomenon thinking collapse: over training, the student's native reasoning behavior erodes as the exploratory tokens that carry it, words like wait, maybe, and alternatively, become progressively less frequent. The mechanism is local rather than global. The damage concentrates at high-entropy decision forks, the branch points where the student is genuinely uncertain and would normally deliberate. Exactly there, the student's top-1 token is often an exploratory marker while the answer-conditioned teacher's top-1 token is not, so the mismatch produces a strong gradient that suppresses the very tokens that make reasoning work. The result is a model whose native reasoning behavior is measurably suppressed, and downstream reasoning accuracy falls in step with it.
probably the best blog i have read for some time
viewing SFT, RL, and OPD as different ways of reshaping a model's distribution makes their tradeoffs super intuitive.
- SFT pulls toward a fixed external target
- RL moves along the reward gradient on on-policy samples
- OPD sits in between, using a teacher signal but on student-generated data, which is why it inherits RL's anti-forgetting properties even when the teacher itself was an overtrained SFT model.
the post is heavily grounded in recent literature and uses the distributional perspective as a unifying bridge across all three paradigms, i really like the point it argues the load-bearing ingredient is on-policy data and OPD's convergence to RL-like outcomes is the strongest evidence
We've solved another piece of the generalist gaming agent puzzle!
Mario requires different skills than Pokemon: reactive navigation, spatial reasoning, and safe exploration rather than long-term memory and zero-sum reasoning.
Check out our new paper on finetuning VLMs to beat Mario via PPO. ๐
Simply adding Gaussian noise to LLMs (one stepโno iterations, no learning rate, no gradients) and ensembling them can achieve performance comparable to or even better than standard GRPO/PPO on math reasoning, coding, writing, and chemistry tasks. We call this algorithm RandOpt.
To verify that this is not limited to specific models, we tested it on Qwen, Llama, OLMo3, and VLMs.
What's behind this? We find that in the Gaussian search neighborhood around pretrained LLMs, diverse task experts are densely distributed โ a regime we term Neural Thickets.
Paper: https://t.co/rFJz2kVEOA
Code: https://t.co/HAmonfpXIA
Website: https://t.co/QZ6AMIsKCw
The idea of rotating attention by 90ยฐ is sooooooo cool (credits to @Jianlin_S 's insights), and it surprisingly works.
We (w/ the amazing @nathan) are so excited about thisโ been working on the paper for months and couldn't stop.
Go give it a try. It's a drop-in replacement for standard residuals, born in 2015.
really like the figs btw :-)
I recently gave a talk at the AI@MIT reading group on our NeurIPS 2025 paper: https://t.co/cXDuIDkkks
We identify the neural mechanism behind attention sinks and propose a training-free mitigation.
Video: https://t.co/TrprZSRgY6
Slides: https://t.co/3L1Mud9c2M
๐งตNew paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck"
The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining ๐
FlashAttention is widely used to accelerate Transformers, already making attention 4-8x faster, but has yet to take advantage of modern GPUs. Weโre releasing FlashAttention-3: 1.5-2x faster on FP16, up to 740 TFLOPS on H100 (75% util), and FP8 gets close to 1.2 PFLOPS!
1/
Test-time errors aren't dead ends. They're training data โ for both our world model and decision model.
Yet embodied systems are stuck with: Fixed world models; Repeated mistakes; Zero growth.
We broke the cycle by introducing Reflective Test-Time Planning โ embodied agents that reflect like human reflective practitioners.
๐ง Reflection-in-Action: We engage in internal simulation, questioning whether our planned approach will actually work given what we currently understand โ before committing to any action.
๐ Reflection-on-Action: We use actual outcomes to reshape both our beliefs about the environment and our strategies for acting within it โ updating the world model and decision policy in real time.
โช Retro-Reflection: Rewind past decisions with hindsight โ so long-horizon failures get the credit assignment they deserve.
๐ https://t.co/CAtCYOIsiWโจ๐ป https://t.co/GDSYrMagJG
Weโre excited to introduce Doc-to-LoRA and Text-to-LoRA, two related research exploring how to make LLM customization faster and more accessible.
https://t.co/ApVzVsBuv1
By training a Hypernetwork to generate LoRA adapters on the fly, these methods allow models to instantly internalize new information or adapt to new tasks.
Biological systems naturally rely on two key cognitive abilities: durable long-term memory to store facts, and rapid adaptation to handle new tasks given limited sensory cues. While modern LLMs are highly capable, they still lack this flexibility. Traditionally, adding long-term memory or adapting an LLM to a specific downstream task requires an expensive and time-consuming model update, such as fine-tuning or context distillation, or relies on memory-intensive long prompts.
To bypass these limitations, our work focuses on the concept of cost amortization. We pay the meta-training cost once to train a hypernetwork capable of producing tasks or document specific LoRAs on demand. This turns what used to be a heavy engineering pipeline into a single, inexpensive forward pass. Instead of performing per-task optimization, the hypernetwork meta-learns update rules to instantly modify an LLM given a new task description or a long document.
In our experiments, Text-to-LoRA successfully specializes models to unseen tasks using just a natural language description. Building on this, Doc-to-LoRA is able to internalize factual documents. On a needle-in-a-haystack task, Doc-to-LoRA achieves near-perfect accuracy on instances five times longer than the base model's context window. It can even generalize to transfer visual information from a vision-language model into a text-only LLM, allowing it to classify images purely through internalized weights.
Importantly, both methods run with sub-second latency, enabling rapid experimentation while avoiding the overhead of traditional model updates. This approach is a step towards lowering the technical barriers of model customization, allowing end-users to specialize foundation models via simple text inputs. We have released our code and papers for the community to explore.
Doc-to-LoRA
Paper: https://t.co/87xEEpf0GN
Code: https://t.co/zBfQi2L9LW
Text-to-LoRA
Paper: https://t.co/emLRZ4Vdvo
Code: https://t.co/b9mrdoWWRB
An ADMM-based optimizer that outperforms state-of-the-art methods across vision models, LLMs, and GANs
Nearly every deep learning model trained today relies on some variant of SGD. Adam, AdamW, Muon, Shampooโthey differ in how they estimate gradient moments or precondition updates, but they share the same theoretical baggage: bounded gradient assumptions, bounded variance, unbiased gradient estimation, and the expectation that training data are IID. When data are heterogeneousโthe norm in federated and distributed settingsโthese assumptions break down and convergence guarantees evaporate.
Shenglong Zhou and coauthors take a fundamentally different path. They build PISA, a preconditioned inexact stochastic ADMM framework that decomposes the training objective into parallelizable subproblems linked through Lagrange multipliers, then solves each inexactly using stochastic gradients and adaptive preconditioning matrices. PISA converges at a linear rate under a single assumptionโLipschitz continuity of the gradient on a bounded regionโwithout requiring bounded variance, bounded gradients, or IID data. Among all stochastic optimizers surveyed, only PISA achieves this combination.
The framework supports pluggable preconditioners, yielding two practical variants: SISA (second-moment preconditioning, analogous to Adam-style adaptive rates) and NSISA (NewtonโSchulz orthogonalized momentum, inspired by Muon). Both retain the same convergence guarantees.
The empirical breadth is notable. Under extreme label skew in federated learning (each client holding one class), SISA reaches 95% on MNIST where FedAvg, FedProx, and Scaffold plateau around 54%. On CIFAR-10, SISA matches or exceeds ten optimizers across ResNet-34, VGG-11, and DenseNet-121. For LLM training, NSISA opens an increasing validation loss gap over Adam, Muon, Shampoo, SOAP, and Adam-mini as model size grows from GPT2-Nano to GPT2-XL, with clear wall-clock advantages at the largest scale. On GAN trainingโnotoriously unstableโSISA achieves the lowest FID scores on both WGAN and WGAN-GP.
What makes this compelling beyond any single benchmark is the theoretical clarity. Most optimizer comparisons are purely empirical; here the convergence theory explains why the algorithm handles heterogeneous data gracefullyโthe ADMM decomposition naturally accommodates non-IID batches without variance reduction over full datasets. The optimization framework itself, not just the learning rate schedule or momentum scheme, can be a decisive design choice for robust training across architectures and data distributions.
Paper: https://t.co/5vTe7kufSR
Finished reading, here is the summary:
SFT before RL usually improves performance or at least saves time/compute. However, when SFT isn't the last stage, measuring a "good" or "successful" SFT becomes non-obvious.
Instead of optimizing for best performance on downstream tasks, SFT should prepare the model for the subsequent RL stage.
This paper shows that strong SFT performance is indeed a bad indicator for success in the subsequent RL stage, with stronger SFT checkpoints often underperforming weaker ones after RL.
The authors introduce a simple weighting scheme during SFT based on the similarity between the behavior policy (the one that generated the SFT data) and the model policy.
The key insight: if the continuation after a token is implausible under the current model, learning from that token provides little useful signal, since RL will sample from its own distribution and rarely visit those trajectories anyway. PEAR down-weights such tokens.
This weighting can operate at the sequence, block, or token level, trading granularity for stability.
They successfully demonstrate that this weighting helps. Interestingly, their experiments use a Qwen model for SFT data generation while training other Qwen models. I was surprised by this, since I expected those policies to already be well aligned.
The only drawback is that the method requires access to the SFT data generation policy. This is only available for synthetic data, and even then, it incurs a non-negligible cost.
However, I like the idea and think optimizing SFT with respect to subsequent RL stages is a good motivation.
Great paper by @dylan_works_ et al.!
Don't use weight decay.
Just normalize everything you see (updates + parameters).
Works on top of your favorite optimizer (e.g., Muon).
Result: 33% speedup + better hyperparameter transfer.