AI will create the most value in three foundational worlds:
1. Information: LLMs and code
2. Physical: embodied intelligence and world models
3. Biological: models that understand, predict, and generate the language of life
Each is a multi-trillion-dollar frontier. I think the third has barely begun, and deserves a lot more attention.
Mathematics is exceptionally compatible with AI because search and reasoning can scale with compute + formal verification tools like Lean. The direction this is all heading is becoming quite difficult to ignore.
We may be watching the beginning of the end of mathematics as an exclusively human profession.
Proximal Policy Optimization (PPO) has been the foundational standard for LLM post-training, yet it quietly suffers from a fatal flaw in long-horizon reasoning: exploration collapse.
🧐 We explore the root cause: PPO suffers from a decade-old Geometric Fallacy!
Excited to share our paper published in ICML 2026:
"Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization"
paper: https://t.co/fgxDMeJPtr
blog: https://t.co/21lhLP2n43
work with: @hello_gensi@ericguoxy@yaqinzhang@haozhou_ai
🚨 PPO’s Geometric Mismatch
🔻 PPO-Clip implicitly adopts a Euclidean metric, resulting in a geometric mismatch with the non-uniform Riemannian policy manifold.
🔻 This mismatch pathologically suffocates trust regions for low-probability tokens while overly expanding high-probability ones, polarizing the policy and triggering exploration collapse!
🚀 Introducing RIPO (Riemannian Isometric Policy Optimization)
🔹 Method: RIPO dynamically aligns isometric trust regions with local geometry, allowing larger updates for low-probability actions while tightening updates for high-probability ones to achieve a balanced exploration-exploitation trade-off.
🔹 Bonus: Geometric isometry mathematically induces statistical homoscedasticity, balancing bias-variance trade-off for fast, stable training!
📌 Key Results
🔹 35% Avg Gain over GRPO:
RIPO consistently refresh SOTA across 7 competition-level mathematical reasoning benchmarks.
🔹 5x Token Efficiency:
RIPO reaches GRPO's 200-step performance in just 40 steps with highly stable gradient norms.
🔹 Sustained Exploration:
RIPO maintains policy entropy at a healthy level throughout late-stage training, preserving policy diversity.
🔹 Unlocks Pass@K Scaling Ceiling:
On HMMT25, RIPO outperforms GRPO by a massive 50% of Pass@128, breaking the intrinsic ceiling of base models!
🌟 Summary
RIPO is simple, principled, and highly effective. By replacing PPO's heuristic Euclidean clip with a dynamic, geometrically aligned clip, it offers a new foundation for RL scaling in LLM reasoning.
As the corresponding author, I'd like to provide an important update regarding CUDA Agent:
After our March arXiv preprint, the Chinese CUDA/AI community identified a small number of reward hacking behaviors in the benchmarking of our CUDA Agent (see Zhihu post below). We have since taken the time to carefully fix the reward hacking issues, and are now excited to share a new updated and improved CUDA Agent! We will be releasing the updated paper soon.
Despite the unexpected challenge, this experience shows the importance of attention to detail in agentic RL. You end up discovering all kinds of interesting and quirky reward hacks when you go digging in the weeds. It also demonstrates that investing time and effort into building a set of robust evals should always be one of the top priorities of agentic RL.
More updates on the paper will be provided in the coming weeks.
@gu_yuxian@thinkymachines Plus, I'm kinda curious about the contribution of this paper. I assume that the Qwen team already has some wonderful distillation techniques that help them to create wonderful Qwen-0.6b/1b/3b... small language models?
@gu_yuxian@thinkymachines +1, I fail to find great differences between TML's on-policy distillation and MiniLLM. To me, on-policy distillation is more like a modified distillation method, and distillation methods were actually studied extensively in 2023 and 2024.