🚨🚨 New paper on flow-matching value functions
Last year, we showed training RL value functions with a flow-matching loss achieved SOTA results.
But why does it work? And what could it possibly tell us about other things that have nothing to do with VFs or even RL?
Short answer: iterative compute used correctly can address feature plasticity in continual learning! 🧵⬇️
@yeonsumia https://t.co/IyuNHsahlY do you think this should be compared here, i feel like TRL is just a DUAL of quasimetric RL but not exactly as the networks are not sticking to the quasimetric assumption. Exploring mixing TD and waypoint and endpoint sampler is maybe the next step?
Introducing PC-ALM, a local-learning alternative to backpropagation.
Our method trains 1000-layer neural nets using only local dynamics, and without backprop.
Blog: https://t.co/bBGCgalqKW
Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop?
We look for inspiration in two related fields: distributed optimization and NeuroAI.
In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors.
This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers.
We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI.
We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors.
We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn.
Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics.
PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs.
Paper: https://t.co/doSZ8mzoyK
Code: https://t.co/rxEDIszVKD
Long-horizon values depend on shorter-horizon values.
But what if those shorter-horizon values are still wrong?
We introduce DCRL: a recursive divide-and-conquer approach that learns values from short to long, substantially mitigating value-error accumulation over long horizons compared with standard GCRL methods. 🧵
Dr. Pratosh at IISc Bengaluru tells his students that rapid advances in AI are commoditizing intellectual labour.
He raises an unsettling question: if companies stop recruiting on campus in the next 5–10 years—even at India’s top universities—what will the purpose of a university education be ?
A must watch video for everyone in tech
Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision
very high benchmarks (beating K3), insane efficiency, and as always amazing tech report
this is probably the most novel arch i've seen in a while, pretty insane
Is RL optimizing the right objective? 🤔
Should we maximize mean reward? Best-of-k? Which k?
Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲
We introduce Tail-Likelihood Reinforcement Learning (TailRL).
It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs.
🧵 1/n
Transformer Transformer will be presented at #CoRL2026 🚀 See y'all in Austin!
Website: https://t.co/s7eimKcHbb
Video: https://t.co/YGQXUSZWLY
Paper: https://t.co/Biyu0r1J7N
Self-supervised RL usually models one action at a time. We extended contrastive RL over action chunks instead, and found large gains: +31.7% offline, +93.1% online. Many explanations exist for why chunking helps. We find a new one: it improves the critic's representations.
🧬Fully Understand Evolution Strategies (ES) for LLM Reasoning
ES can post-train LLMs without backprop. Our study finds that ES:
🔍 Explores higher Pass@K than GRPO
💥 Still sparse functional updates, with no catastrophic forgetting.
⚙️ Requires fewer samples as LLMs get larger
We summarize the paper through 3 research questions 👇
1️⃣ Does ES exhibit the same post-training characteristics as GRPO?
ES maintains broader reasoning coverage!
Its population-based parameter-space exploration induces diverse reasoning behaviors: ES improves Pass@1 while achieving stronger large-K Pass@K, without the same entropy collapse observed in GRPO.
The two are also complementary. 🔀 Sequential compositions—ES→GRPO and GRPO→ES—combine GRPO’s strength in Pass@1 with ES’s gains in Pass@K, yielding better Pass@1–Pass@K trade-offs.
➡️ ES preserves broader coverage; combining it with GRPO can capture both strengths.
2️⃣ Does ES necessarily cause catastrophic forgetting?
No. Although ES induces substantial whole-model parameter drift, its task-relevant effects are concentrated in a small subset of larger-magnitude updates.
Most parameter changes contribute little after perturbation cancellation, revealing surprising functional sparsity. Held-out capabilities also remain largely preserved under appropriate training settings.
➡️ Large parameter movement does not necessarily imply widespread functional change—or catastrophic forgetting.
3️⃣ What hyperparameter settings and estimators make ES effective and scalable?
We find three practical takeaways:
✅ Z-score reward normalization is a key ingredient for effective ES training.
✅ Larger pretrained models require smaller population sizes for effective optimization.
✅ For discrete reasoning rewards, the two-point estimator commonly favored in zeroth-order SFT provides no advantage.
💡 Takeaway: ES is not merely a memory-efficient, gradient-free alternative to GRPO. Its broader reasoning coverage, functionally sparse updates, and distinct scaling behavior position it as a distinct paradigm for LLM reasoning post-training.
📄 arXiv:2608.27351
Generative models can’t discover what they can’t reach.
We’re excited to introduce ActFlow: a continued pre-training scheme that actively expands the valid design space reachable by flow and diffusion models. We call this generable set expansion — a new learning principle for out-of-distribution generative modeling, and a step toward evolvable search spaces for scientific discovery. (1/5)
Behavioral cloning mystery
https://t.co/VqxzvcfGSx
I wrote a new blog post about "mysteries" in behavioral cloning that appear with real-world robot data (e.g., overfitting is "good"). I also tried to demystify them and shared my thoughts!
Excited to introduce our work, Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy, which has been accepted to ICML 2026!
By leveraging flow-consistent values, we resolve the critical trade-off between expressivity and stability in Flow-based Reinforcement Learning.
Joint work at KAIST w/ @bkjeon1211 , @SeonghyeonYe , @kimin_le2 , @seo_minjoon .
Paper: https://t.co/J5FR6ac9OF
Code: https://t.co/1tObfGord1
Project Page: https://t.co/BF0JWAKSKk
@klindt_david I tried training a deepkoopman space vs le-wm Gaussian distribution space and the koopman space fits better but planning isn’t possible even when space is linear as singularities can be anywhere and I couldn’t find a linear way to model the usable cspace
@klindt_david What if the underlying data space doesn’t have a similar topology, this breaks imo, this said neural nets learn a high dimension approximation, so it’s possible that it can have singularities to the edge of the distribution.
🧠We introduce "Generative Recursive Reasoning"!
Recursive Reasoning Models like HRM, TRM, and Looped Transformers are deterministic — same input, same reasoning, every time. They collapse the entire space of plausible reasoning paths into a single attractor.
Our model GRAM (Generative Recursive reAsoning Models) turns recursion itself into a stochastic latent trajectory. Multiple hypotheses, alternative solution strategies, and inference-time scaling not just by depth, but by width — parallel trajectory sampling.
And here's the kicker: the same formulation that gives us conditional reasoning p(y|x) also makes GRAM a general generative model p(x).
With only 10M params:
• Sudoku-Extreme: 97.0% (TRM 87.4%)
• ARC-AGI-1: 52.0%
• ARC-AGI-2: 11.1%
• N-Queens coverage: 90%+
📄 Paper: https://t.co/JC7EyXYc9Y
🌐 Project page: https://t.co/LRT1dQiWLZ
w/
Junyeob Baek @JunyeobB (KAIST),
Mingyu Jo @pyross0000 (KAIST),
Minsu Kim @minsuuukim (KAIST & Mila),
Mengye Ren @mengyer (NYU),
Yoshua Bengio @Yoshua_Bengio (Mila),
Sungjin Ahn @SungjinAhn_ (KAIST)
The co-inventor of Looped Transformers defended her PhD thesis yesterday and is heading to an incredible new role soon :) congratulations @AngelikiGiannou 🥳 🎉🎈