Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse.
๐ง We explore the root cause: PPO suffers from a decade-old Geometric Fallacy!
Excited to share our paper published in ICML 2026:
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
paper: https://t.co/fgxDMeJPtr
blog: https://t.co/21lhLP2n43
work with: @hello_gensi@ericguoxy@yaqinzhang@haozhou_ai
๐จ PPOโs Geometric Mismatch
๐ป PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold.
๐ป This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse!
๐ Introducing RIPO (Riemannian Isometric Policy Optimization)
๐น Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off.
๐น Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training!
๐ Key Results
๐น 35% Avg Gain over GRPO:
RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks.
๐น 5x Token Efficiency:
RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms.
๐น Sustained Exploration:
RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity.
๐น Unlocks Pass@K Scaling Ceiling:
On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models!
๐ Summary
RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
PPO-Clip is the default for LLM RL, but it accidentally kills exploration by treating all probability ratio changes as equally large.
This paper shows the bug is geometric. Rare tokens and common tokens live in very different parts of the policy manifold, so fixed clipping under-updates rare but useful reasoning paths and over-updates already dominant ones.
They fix this with probability-aware Riemannian clipping, giving rare actions more room and common actions less.
This gives smoother RL, less exploration collapse, and up to 60% better AIME24 performance over GRPO.
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse.
๐ง We explore the root cause: PPO suffers from a decade-old Geometric Fallacy!
Excited to share our paper published in ICML 2026:
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
paper: https://t.co/fgxDMeJPtr
blog: https://t.co/21lhLP2n43
work with: @hello_gensi@ericguoxy@yaqinzhang@haozhou_ai
๐จ PPOโs Geometric Mismatch
๐ป PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold.
๐ป This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse!
๐ Introducing RIPO (Riemannian Isometric Policy Optimization)
๐น Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off.
๐น Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training!
๐ Key Results
๐น 35% Avg Gain over GRPO:
RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks.
๐น 5x Token Efficiency:
RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms.
๐น Sustained Exploration:
RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity.
๐น Unlocks Pass@K Scaling Ceiling:
On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models!
๐ Summary
RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
@xinbinjian Yeah, that's right.
And RIPO's computational complexity is negligible compared to the extremely large computation of LLM's rollout and gradient-update
Thanks for highlighting our work!
PPOโs Euclidean Clip is geometrically mismatched with the Riemannian policy manifold, causing exploration collapse.
RIPO aligns trust regions with local geometry, achieving superior LLM RL performance!
Happy to discuss!
https://t.co/4lGdwjF8vs
@sheriyuo Thanks for highlighting our work!
PPOโs Euclidean Clip is geometrically mismatched with the Riemannian policy manifold, causing exploration collapse.
RIPO aligns trust regions with local geometry, achieving superior LLM RL performance!
Happy to discuss!
https://t.co/4lGdwjF8vs
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse.
๐ง We explore the root cause: PPO suffers from a decade-old Geometric Fallacy!
Excited to share our paper published in ICML 2026:
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
paper: https://t.co/fgxDMeJPtr
blog: https://t.co/21lhLP2n43
work with: @hello_gensi@ericguoxy@yaqinzhang@haozhou_ai
๐จ PPOโs Geometric Mismatch
๐ป PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold.
๐ป This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse!
๐ Introducing RIPO (Riemannian Isometric Policy Optimization)
๐น Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off.
๐น Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training!
๐ Key Results
๐น 35% Avg Gain over GRPO:
RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks.
๐น 5x Token Efficiency:
RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms.
๐น Sustained Exploration:
RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity.
๐น Unlocks Pass@K Scaling Ceiling:
On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models!
๐ Summary
RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.