@COLM_conf@SusanMurphylab1@YaelBatat If you're in the Bay (perhaps for COLM or Tech Week), and you're interested in Computational Social Science, LLMs, and Causal Inference, come check out our talk!
Krutch Auditorium @UCBerkeley , 3:15 PM, Panel 3 (hosted by @dbamman )
On October 5th @trsaenger and I will give an oral presentation at @textasdata at @UCBerkeley.
We’re incredibly excited to present "APPEAL: Attributing Persuasive Power to Expert-Specified Actions in LLMs"
This is joint work with @GiladMorad12, @ofraam, and @b_m_stewart
Thank you to my advisors, @ofraam and @amir_feder, for supporting me on my trip and continuing to guide my research despite the distance and time differences.
Thank you to @b_m_stewart for his guidance and support. I couldn’t have asked for a more dedicated and brilliant mentor
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
Is RL actually making your LLM better?
Gains from RL are mostly on easy questions🤯
We're calling this the Matthew Effect for RL on LLMs. We then leverage async RL to solve harder problems by Never Giving Up!
paper https://t.co/1C2xjunbWc blog https://t.co/X4H83wKgXl 🧵👇
Is RL optimizing the right objective? 🤔
Should we maximize mean reward? Best-of-k? Which k?
Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲
We introduce Tail-Likelihood Reinforcement Learning (TailRL).
It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs.
🧵 1/n
I recently lost access to my Twitter/X account, @ZacharyBamberg1 . I noticed today that my old account published a post that I did not write. If you get a DM from this account, note that it is NOT ME. I also urge you to report this old account as hacked if you have the time.
I just published my complete guide to reinforcement learning for LLMs. It's a single, standalone resource for understanding RL from first principles to frontier research.
Read it here: https://t.co/iUMAWUrkQ1
The topics covered are as follows…
RL Fundamentals:
- A general framework for RL (policy, states, actions, transition functions, environment, rewards, trajectories, etc.)
- Core concepts of RL (returns, discounting, action / state / trajectory probabilities, value / advantage functions).
- Formulating and estimating the learning objective for RL.
Basic of RL for LLMs:
- Token-level (MDP) vs. completion-level (bandit) formulations.
- Outcome versus process rewards.
- Value estimation via value models / critics.
- KL divergence, RL setups (RLHF, RLVR, infra, etc.), importance sampling.
Basic Policy Gradients:
- Deriving the RL objective / vanilla policy gradient (VPG) from first principles.
- Different policy-gradient formulations (trajectory returns, reward-to-go, baselines, value functions, and advantages).
- Why policy gradient estimates have high variance (and how baselines help).
- Implementation of VPG.
REINFORCE and its variants:
- REINFORCE as a Monte Carlo implementation of the VPG.
- Completion-level vs. token-level implementations of REINFORCE.
- Reward baselines and KL-regularized rewards.
- RLOO, REINFORCE++ and other critic-free policy gradient algorithms.
Actor-Critic Methods:
- Training value models / critics alongside the policy to estimate advantages.
- Trust Region Policy Optimization (TRPO) and KL-constrained policy updates.
- How PPO simplifies TRPO with a clipped surrogate objective.
- Full PPO details (clipping, advantage estimations, value loss, KL, policy ratios, multi-epoch optimization, etc.) + implementation.
- Generalized Advantage Estimation (GAE) + implementation.
- The memory and computational overhead of actor-critic algorithms.
GRPO:
- Group Relative Policy Optimization (GRPO) as a simpler alternative to PPO.
- Group-relative advantage estimates + implementation.
- Why removing the critic is helpful.
- Limitations of vanilla GRPO algorithm.
- Dr. GRPO, DAPO, GSPO, and other recent improvements.
Recent research topics:
- Online versus offline RL.
- Rubric-based RL.
- RL scaling laws.
- Agentic RL + world modeling.
This post is a synthesis of my writing / learning on RL for over a year. I hope it’s helpful!