🤖Introducing OptimalThinkingBench 🤖
📝: https://t.co/cffhCY4eQw
- Thinking LLMs use a lot of tokens & overthink; non-thinking LLMs underthink & underperform.
- We introduce a benchmark which scores models in the quest to find the best mix.
- OptimalThinkingBench reports the F1 score mixing OverThinkingBench (simple queries in 72 domains) & UnderThinkingBench (11 challenging reasoning tasks).
- We evaluate 33 different SOTA models & find improvements are needed!
🧵1/5
Are RL agents truly learning to reason, or just finding lucky shortcuts? 🤔
Introducing RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards — a novel framework that rewards not just outcomes, but the quality of reasoning itself, creating more robust and generalizable agents.
1️⃣ We identify "inefficient exploration" in standard RL: agents achieve success through flawed reasoning paths (e.g., repetitive actions, illogical steps), leading to brittle policies that fail on new tasks.
2️⃣ RLVMR provides dense, process-level rewards for verifiable meta-reasoning behaviors:
• 🎯 Planning: Reward strategic thinking
• 🔍 Exploration: Reward discovering new states
• 💭 Reflection: Reward error correction
3️⃣ Results on ALFWorld & ScienceWorld:
• 🏆 New SOTA: 83.6% success on hardest unseen tasks (7B model)
• 📉 Significant reduction in repetitive actions
• 🚀 Enhanced generalization to novel scenarios
🧑💻 Code: https://t.co/4eKmiCZa8K
📃 Paper: https://t.co/2mthD4aSaR
Pruning is an effective way to speed up LLM inference. However, most existing methods are static. In our #ICML 2025 paper, we propose a novel dynamic pruning method, which achieves comparable or even better performance than the base model despite a 40% reduction in parameters.
We've taught LLMs math and code with RLVR. But can we teach them empathy? 🤖❤️
Introducing Reinforcement Learning with Verifiable Emotion Rewards (RLVER), the first RLVR framework that enhances LLMs' empathy from a simulated user .
❤️ Feelings → Numbers: A psychologically-grounded user simulator (SAGE) delivers transparent, deterministic, audit-ready emotion scores after every dialogue, turning "feelings" into RL signals.
🚀 Results: an open-source 7B model’s Sentient-Benchmark score leaps from 13.3 ➡️ 79.2, rivaling proprietary models 10× its size while preserving coding & math skills.
🧐 Training Insights
1⃣ Thinking vs. non-thinking routes diverge: thinking lifts empathy/insight; non-thinking favors action.
2⃣ GRPO = steadier gains, PPO = higher peaks.
3⃣ Moderately challenging environments beat overly hard ones for EQ growth.
🤝 We’re open-sourcing code, checkpoints, and scripts to accelerate research into emotionally intelligent AI!
🧑💻 Code & Model: https://t.co/6Pi46rVZVC
📃 Paper:
https://t.co/8RrBsKkaww
🚀 Atomic-to-Compositional Generalization for Mobile Agents
🧠 A new benchmark & scheduling system to push the limits of mobile agens. 📄 Paper: https://t.co/Z4X3ReGgFy 🌐 Website: https://t.co/gy7zy2i0JN
🧵1/n
🚀 Check out our paper: WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model, from Tencent AI Lab!.
We present a world model-driven framework for self-improving web agents, addressing critical challenges in self-training—such as limited exploration and performance plateaus.
🔍 Key Innovations:
- Co-Evolving World Model:
The world model is implemented as an LLM that predicts the next webpage state (observation) given the current state and a planned action.
In addition to fine-tuning the agent's policy model using self-generated trajectories, the same data is repurposed to train the world model.
- World Model as Web Server
Training Phase: sample pseudo trajectories by replacing the real web server with the world model.
Inference Phase: simulate the outcome of candidate actions with 1-2 step look-ahead planning, to help better select actions.
📊 Results on Real-World Tasks (Mind2Web-Live & WebVoyager):
✅ ~10% higher success rate vs. pure self-training.
✅ significantly fewer environment interactions—efficient yet powerful!
arxiv: https://t.co/7tMVy4iI0J
code: https://t.co/wzutAx8zNR
🚀 Introducing ROLL: An Efficient and User-Friendly RL Training Framework for Large-Scale Learning!
🔥 Efficient, Scalable & Flexible – Train 200B+ models with 5D parallelism (TP/PP/CP/EP/DP), seamless vLLM/SGLang switching, async multi-env rollout for maximum RL throughput!
⏰ We introduce Reinforcement Pre-Training (RPT🍒)
— reframing next-token prediction as a reasoning task using RLVR
✅ General-purpose reasoning
📑 Scalable RL on web corpus
📈 Stronger pre-training + RLVR results
🚀 Allow allocate more compute on specific tokens
This year, there have been various pieces of evidence that AI agents are starting to be able to conduct scientific research and produce papers end-to-end, at a level where some of these generated papers were already accepted by top-tier conferences/workshops.
Intology’s AI-generated paper getting accepted by ACL is one example: https://t.co/NcWHmD9f08
In fact, three independent teams submitted AI-generated papers to ICLR’25 workshops and got some of them accepted:
Sakana: https://t.co/f3A2ihIEO7
AutoScience: https://t.co/P9qqKOTvPh
Intology: https://t.co/xTaB4DXD09
We will probably continue to see AI-generated submissions at various conferences/journals, and maybe more will be accepted. But what does it tell us? Are LLMs already better than us human researchers? Here are my two cents:
1. I do believe LLMs will be able to automate parts of our research pipeline, and this could be a good thing.
It seems that directly building agent scaffolds on top of current LLMs can already get you some non-trivial results (all of the above systems use agent scaffolds on top of commercial LLM APIs).
Not to mention that there has been a thrust of recent efforts on building research agent environments with objective rewards (e.g., benchmark performances): RE-Bench, MLE-Bench, PaperBench, MLGym, MLE-Dojo, just to name a few.
Just like the rapid development of coding agents, I believe research execution agents will keep getting better as we develop better training algorithms (e.g., RL) on these environments and as we have more capable base models to work with.
And I think this is good for us because we can offload some of the tedious implementation work to these agents to speed up our research once they start to get reliable.
2. Producing research papers end-to-end and submitting AI-generated papers to conferences, on the other hand, is bad in many ways.
For these research agent developers, submitting AI papers to peer review is a terrible way to do evaluation. Whether you like it or not, you have to admit that conference reviewing can be quite noisy. And I just don’t see how you can make any statistically meaningful conclusions from the result of this one single paper acceptance in this noisy reviewing process, not to mention all the extra burden you are imposing on our (already partially broken) conference reviewing system.
For the broader scientific community, I just don’t think the purpose of building AI Scientists is to produce thousands of mediocre research papers that may have a chance of passing the conference thresholds. Getting a paper accepted is one thing; actually pushing the frontier of human knowledge is a completely different matter. Our community will be doomed if everyone (including the AI Scientists) treats conference acceptances as the only reward function. For all of us working on AI Scientists, I think the right goal we should be pursuing, is how AI Scientists can help us make scientific breakthroughs that we humans alone wouldn’t be able to achieve, rather than hacking flawed reward functions.
This is not an easy post to write because I know all of these teams working on AI Scientists personally, and these are some really smart people that I deeply respect. If there’s only one message to take away from this post, I hope all of us working on AI Scientists can work towards the right objectives and try to make this world a better place. ❤️
🚨 Announcing DeepTheorem: Revolutionizing LLM Mathematical Reasoning! 🚀
𝕋𝕃𝔻ℝ:
- 🌟 Learning by exploration is the most important rationale that recent RL-zero training teaches us since self-exploration significantly boosts the utilization of LLM pre-training knowledge;
- 🧐 Since LLM is pre-trained with massive knowledge of mathematical theorems, can LLM learns theorem proving by self-exploration?
- 🤯 We show that using our high-quality deep theorem dataset with online RL learning is sufficient to activate LLM's theorem-proving ability. Our 7B model can outperform even advanced models like **Gemini** and **Claude 3.5**! More importantly, we don't need any theorem proof annotation, all you need is the truth value of the theorem itself.
- 📄Come and check our paper:
Arxiv: https://t.co/2686VIMcre
Huggingface: https://t.co/RhLUkFaJK0
@z4y5f3 Thanks for sharing your findings. We used temperature-based sampling and reported pass@1 (n=16), so it is normal for the results to be inconsistent with greedy decoding. In addition, vllm's greedy decoding is indeed unstable, which means it is likely not a true greedy decoding.
This post really resonates with me—evaluation is so crucial yet so challenging. Let me share some experiences from our DeepMath-103K project:
We've noticed multiple works reporting results in rather "tricky" ways. For instance, using greedy decoding + selecting the checkpoint with highest AIME24 score. As we know, benchmarks like AIME24 (with only 30 problems) show high variance under greedy decoding (just changing the training seed or step can cause significant fluctuations).
This leads to two issues:
1. Performance drops noticeably (~10 points) when switching to temperature-based sampling
2. While AIME24 results look great, other benchmarks (AIME25, OlympiadBench, etc.) show almost no improvement.
BTW, working on DeepMath made me realize how thorough our evaluation approach was:
1. We rigorously decontaminated data based on semantics across 10+ benchmarks
2.Aligned prompt templates with those used in baseline models' training
3. When our baseline reproductions fell significantly short of reported results, we either contacted authors or investigated GitHub issues (often finding others couldn't reproduce either)
For reference:
Our evaluation shows Qwen-2.5-7B achieves 54.8 on MATH500, higher than Spurious Rewards (41.6) and Entropy Minimization (43.8), but lower than Sober Look (64.6). Interestingly, Qwen's TR only reports 49.8 on MATH.
Our Qwen-2.5-Math-7B reaches 46.9, higher than LRM-Self-train (~42), but still lower than Sober Look (64.3). Qwen's TR reports 55.4 on MATH (4-shot).
R1-Distill-Qwen-1.5B reaches 84.7, higher than 1-shot RLVR(71.9), and match Sober Look (84.9).
https://t.co/yYAQKCN3Of
Confused about recent LLM RL results where models improve without any ground-truth signal? We were too. Until we looked at the reported numbers of the Pre-RL models and realized they were serverely underreported across papers. We compiled discrepancies in a blog below🧵👇