Full paper: https://t.co/wPaKRvPZzi. If you're at NeurIPS2025, come see us at Exhibit Hall C,D,E - Poster #3407, on Wednesday, Dec 3, 11 AM – 2 PM PST!
RL methods like PPO reduce the probability of sampled low-reward sequences. With training, these outputs become rare, with weaker gradient signal. We can better reduce their probabilities using probabilistic inference to aid sampling! (with @aidanmrli@brekelmaniac@RogerGrosse)
Announcing Transluce, a nonprofit research lab building open source, scalable technology for understanding AI systems and steering them in the public interest.
Read a letter from the co-founders Jacob Steinhardt and Sarah Schwettmann:
https://t.co/IUIhBjpYhS
RL is not the only paradigm for language model alignment or controlled generation!
Let’s take a principled probabilistic approach: define a target distribution, perform inference, estimate KLs. Enter our work on twisted SMC for LMs with @brekelmaniac* @AliMakhzani@RogerGrosse
We use our bidirectional SMC bounds to evaluate inference methods (including PPO) in these settings, via forward and reverse KL divergences between the model q(𝐬) and target σ(𝐬) distributions, showing our framework’s ability to provide meaningful quantitative evaluation.
SOTA AI algorithms in general-sum settings often result in worst case collective outcomes. Opponent shaping (eg LOLA https://t.co/KloekN99YB) partially addresses this. We show that prior work is sensitive to policy parameterization and introduce a novel method to solve this issue
Check out the full paper: https://t.co/0qtvS6swFY, come see us at NeurIPS in person or online: https://t.co/9yOipvWFWJ, and/or send comments and questions our way!