What advantage to use, and when? Everyone's proposing new advantage functions for RL with LLMs but nobody knows why they work or fail.
We break this down and build FADE a self-adapting advantage to get +14% on LiveCodeBench v6 in 40% less steps.
Paper: https://t.co/fAX16uYPpt
Cool paper! We can view score centering as:
REINFORCE+R*KL(q||p) since grad_p KL = -E_q
so RL went from KL reg -> no KL -> KL bonus!
Funny enough I tried w/ @TacoCohen to revive KL reg by smoothing pi_ref. Realized now we don't know where to push p, only what to avoid: q's bias.
RL with LLMs is very unstable when training and sampling policies differ. Standard fixes (matching numerics, importance sampling) work around the problem.
We find the root cause of this instability from first principles and propose a way to directly cancel it. Score Centering is competitive and compatible with existing approaches — while simple to implement!
🧵 [1/6]
WaiT for the Signal: Simple Frequency-Aware Flow-Matching.
We show how to natively incorporate a fundamental property of images directly into diffusion models; setting a new pixel-space SOTA on ImageNet, while reducing compute.
📄 https://t.co/5ZJVOcfzQA
Full breakdown below👇
How can we get LLMs to produce fast code?
Naive: easy just reward correctness + speed
The truth: it's hard to measure, hard to reward, hard to train.
Checkout all of @PierreChambon6 complex figures to see how we can get it to work:
But now new problem: RL envs * parametrization then becomes too big
-> we need to explore this space of RL envs in efficient manner, as it would burn too much gpus.
So we screen everything with offline RL simulator: turns out the space is very sparse and lots of envs are bad.
When I showed @PierreChambon6 the curve of the correctness-efficency frontier on code RL, he's like:
"Let's just RL on code run time"!
Turns out that's a huge rabbit hole, but here we are finally, after cleaning about all the mess. Congrats!
🧵How to combine code correctness and efficiency objectives in online RL?
-> we break the correctness-efficiency Pareto frontier!
+125% relative improvement in CWM32B compared to standard RLVR (13.7 ->30.9 pass@1@top30%)
+150% with Qwen32B
Paper: https://t.co/Wdn5vlPxmV
🇰🇷 #ICML2026 Alert! 🇰🇷
Come check out our new work with @mathuvu_ and @ylecun 👀
📍 Starting 2PM - Poster #1606
💡 Representation learning for text done efficiently - clean scaling, 2× retrieval, ~100× less compute.
🚨 Spoiler Alert: BERT does not scale
https://t.co/kh8nDYaYUC Really cute work from @AIatMeta’s CodeGen team in DecompRL, most notably @DecugisJuliette and @FabianGloeckle
You train a code model first to decompose a problem in multiple functions, then implement each of those -> through permutations of the rollouts of the implementations, you can actually size a diverse code training data much easier
They make a specific policy gradient algorithm to grade that process, first training the decomposition then the implementation policy using a logmeanexp function to compute the objective of the multiple rollouts and aggregating over a leave-one-out baseline
Models trained are as good as other RL algorithms… but work better on larger token budget, and most notably have much less emphasis on GPU (less need for rollouts) and much more emphasis on CPU (more need for function verification)
A pretty cool work that makes you think about domain specific training! 😊
I’m at ICML🇰🇷 this week in Seoul !
Feel free to reach out if you want to chat about RL and codegen!
Will also present @DL4Code my work on advantages and at the RLxF workshop work on efficient codegen (led by @PierreChambon6)!
@PengmingWang Yes agreed but here difficulty is defined relative to the current policy. With a fixed dataset, early in training you want to maximize learning signal on solvable hard tasks. As the policy improves, there are less hard tasks while medium tasks provide denser, reliable updates.
What advantage to use, and when? Everyone's proposing new advantage functions for RL with LLMs but nobody knows why they work or fail.
We break this down and build FADE a self-adapting advantage to get +14% on LiveCodeBench v6 in 40% less steps.
Paper: https://t.co/fAX16uYPpt
Joint work with @seano_research and my amazing advisors: @BachFrancis , @syhw , @TacoCohen
We hope this gives the community a way to sort through the advantage literature and opens the door to dynamic allocation of gradient mass. 🚀
Open questions we're excited about:
- how can we extend this framework to noisy verifiers? (eg, with judge models)
- can we get partial credit from failures to make failure-only RL learn?
🧠⌨️ Decode language from brain activity without surgery. 🧠⌨️
Brain2Qwerty V1 is officially published in Nature Neuroscience. Today, we're releasing Brain2Qwerty V2.
We achieve unprecedented performance for a non-invasive MEG setup.
Details below 🧵👇
Amazing work led by @KunhaoZ Turns out your model either optimises speed or code correctness during RL and you can minimise both failures by extrapolating between RL checkpoints! Let's see how we can bring this to training 🧐
🧵 For 2 RL checkpoints trained differently, you can just weight extrapolate them and it works!
Bonus: these extrapolated checkpoints are complementary policies
-> Get exploration and diversity for free
-> Better inference scaling when ensembling
Paper: https://t.co/zU0LH0TOdm