I just published my complete guide to reinforcement learning for LLMs. It's a single, standalone resource for understanding RL from first principles to frontier research.
Read it here: https://t.co/iUMAWUrkQ1
The topics covered are as follows…
RL Fundamentals:
- A general framework for RL (policy, states, actions, transition functions, environment, rewards, trajectories, etc.)
- Core concepts of RL (returns, discounting, action / state / trajectory probabilities, value / advantage functions).
- Formulating and estimating the learning objective for RL.
Basic of RL for LLMs:
- Token-level (MDP) vs. completion-level (bandit) formulations.
- Outcome versus process rewards.
- Value estimation via value models / critics.
- KL divergence, RL setups (RLHF, RLVR, infra, etc.), importance sampling.
Basic Policy Gradients:
- Deriving the RL objective / vanilla policy gradient (VPG) from first principles.
- Different policy-gradient formulations (trajectory returns, reward-to-go, baselines, value functions, and advantages).
- Why policy gradient estimates have high variance (and how baselines help).
- Implementation of VPG.
REINFORCE and its variants:
- REINFORCE as a Monte Carlo implementation of the VPG.
- Completion-level vs. token-level implementations of REINFORCE.
- Reward baselines and KL-regularized rewards.
- RLOO, REINFORCE++ and other critic-free policy gradient algorithms.
Actor-Critic Methods:
- Training value models / critics alongside the policy to estimate advantages.
- Trust Region Policy Optimization (TRPO) and KL-constrained policy updates.
- How PPO simplifies TRPO with a clipped surrogate objective.
- Full PPO details (clipping, advantage estimations, value loss, KL, policy ratios, multi-epoch optimization, etc.) + implementation.
- Generalized Advantage Estimation (GAE) + implementation.
- The memory and computational overhead of actor-critic algorithms.
GRPO:
- Group Relative Policy Optimization (GRPO) as a simpler alternative to PPO.
- Group-relative advantage estimates + implementation.
- Why removing the critic is helpful.
- Limitations of vanilla GRPO algorithm.
- Dr. GRPO, DAPO, GSPO, and other recent improvements.
Recent research topics:
- Online versus offline RL.
- Rubric-based RL.
- RL scaling laws.
- Agentic RL + world modeling.
This post is a synthesis of my writing / learning on RL for over a year. I hope it’s helpful!
This is one of the most useful writeups I have seen on keeping an LLM judge effective in production.
(bookmark it)
Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members on mobile.
They describe the judge as a lifecycle with four phases rather than an artifact that you only validate once.
> Birth defines multiple evaluation criteria and builds curated benchmarks with human labels and rationales.
> Training refines the judge's rubric through Reasoning-Aligned Rubric Tuning, using a meta-judge over reasoning output as the learning signal.
> Deployment puts one judge in two roles, quality gating and reflective generation.
> Monitoring runs continuous human-in-the-loop alignment that detects drift and triggers re-tuning behind a review gate.
A five-week A/B test over tens of millions of members shifted viewing toward previously unwatched content and increased successful browse-to-play sessions against a no-explanation control, with no quality-related takedowns.
Paper: https://t.co/H12znqteUi
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
Getting ready to publish my complete guide to RL for LLMs tomorrow morning. Although the post contains many of my own thoughts / learnings, it is also a synthesis of so many great resources that have been published over the years:
- The RLHF Book (https://t.co/n3aUxqWtgS) by @natolambert
- Reinforcement Learning by Richard S. Sutton and Andrew G. Barto
- Spinning Up in Deep RL (https://t.co/EYcOOllvIy) from OpenAI
- Build an LLM (https://t.co/6HfXuofVwq) and Reasoning Model (https://t.co/l5uS247A04) from Scratch by @rasbt
- Various notes (https://t.co/hpsyAKOnRv) and papers (TRPO, PPO, etc.) from John Schulman
- Policy Gradient Algorithms (https://t.co/caNLakSjo5) by Lilian Weng
- A Vision Researcher’s Guide to RL (https://t.co/25awrjj5Bd) by @YugeTen
- From REINFORCE to Dr. GRPO (https://t.co/M23e3Fl3VZ) by @qingfeng_lan
- Async GRPO in the Wild (https://t.co/A5qeYMwcy4) by @yumo_xu
- Open RL infrastructure like TRL (https://t.co/TGrrnJ574t) and OpenInstruct (https://t.co/L9pObiUb1c)
I highly recommend reading all of them. They’ve truly helped me to learn so much.
A new skyline is emerging across Ahmedabad with the Mumbai–Ahmedabad Bullet Train Project.
2 Bullet Train stations
1 Sabarmati Multimodal transport hub
18 km of viaduct
6 steel bridges
1 river bridge
Together, these structures are transforming Ahmedabad’s landscape and shaping a new era of connectivity for the city.
#BharatKaGarv
@uginm102@shaaka777@kinjeketile@TheMutaD Sorry... I meant to write
"I know that 5% of 20 Trn is more than 8% of 4 Trn."
Also, that 5% for China is real GDP figure. The nominal figure is less than 5% due deflation.
But as I said, keep compounding for 20 years.
@uginm102@shaaka777@kinjeketile@TheMutaD Where was I irrational? Did I make any ridiculous statement like comparing Nigeria and India.
I know the math. I know that 5% of 20 Trn is less than 8% of 4 Trn. But keep doing the compounding for 20 years.
Strap in. 🚀
Ride along with Vikram-1 from the rocket’s POV. Earth slips away, stages separate, the darkness of space opens up, and our blue planet comes into view as orbit is achieved.
🎥 Onboard camera footage.
#MissionAagamanInAction#Vikram1
@shaaka777@uginm102@kinjeketile@TheMutaD But sure! India = Nigeria.
Even if you want to compare Nigeria within Africa and India within Asia, what a ridiculous comparison?
@uginm102@shaaka777@kinjeketile@TheMutaD LOL!
Uganda has a 1/3rd the GDP per capita of India and is not even growing at the same speed as India, but look at the audacity?
@uginm102@shaaka777@kinjeketile@TheMutaD China at 5%?
LOL! It is not growing at 5% even now. It's population is declining. Even with all the AI and export promotion, it is not going to grow at even 5%.
India will continue to grow at 7-8% for next 2 decades. India's current demographics is what China had in late 1990s
@shaaka777@uginm102@kinjeketile@TheMutaD "One thing I can bet on is India will not only not catch China, but China will actually gap them more."
This was the statement earlier in this thread. But people forget that India has been growing faster than China since at least 2016.
@shaaka777@uginm102@kinjeketile@TheMutaD Do you have any GDP growth data to back up your claim of India underperforming?
India has been growing at 7-8% since covid. Of course, it is not 2000's era Chinese level 10-12% growth, but it is still faster than any country in Asia.