Excited to share our paper "Programming over Thinking: Efficient and Robust Multi-Constraint Planning" at #ACL2026! 🎉
Due to visa issues, I'll be presenting virtually at Virtual Presentations 3 on Mon, Jul 6, 12:45–14:15 PDT. Hope to connect 😃
🎊Thrilled to share that my work during my last tenure in Huawei is accepted in #EMNLP2025 (Main).
Paper : https://t.co/afrhBjCpqC
Code : https://t.co/uBrCaOxN40
Dataset : https://t.co/AFYhK7Jbg0
#emnlp#multimodal#vllm
🎊Thrilled to share that my work during my last tenure in Huawei is accepted in #EMNLP2025 (Main).
Paper : https://t.co/afrhBjCpqC
Code : https://t.co/uBrCaOxN40
Dataset : https://t.co/AFYhK7Jbg0
#emnlp#multimodal#vllm
Seems like no one saw this either, scraping arxiv manually seems to be the way. Pretty cool paper on rl for creative writing on Qwen3 32B base, and most interestingly it's one author from the Star Writing Team (haven't heard of them). They seem to have access to the 32B base tho so I guess it's a part of the Qwen team?
🚨Self-Challenging Language Model Agents🚨
📝: https://t.co/skxQDwuNk2
A new paradigm to train LLM agents to use different tools with challenging self-generated data ONLY: Self-challenging agents (SCA) both propose new tasks and solve them, using self-generated verifiers to derive reward for RL training.
Training on self-synthesized tool-use trajectories, SCA significantly boosts the base LLM’s tool-use capabilities, with over 2x improvement on TauBench and M3ToolEval.
🧵1/4
@RulinShao Hi @RulinShao , thank you for your thoughtful reply. I will try running the script. Hopefully there will be also be released checkpoints as standardized baseline.
There are traditionally two types of research: problem-driven research and method-driven research. As we’ve seen with large language models and now AlphaEvolve, it should be very clear now that total method-driven research is a huge opportunity.
Problem-driven research is nice because you have a consistent and specific goal. The goal is usually virtuous, so it feels good to have a mission and identity. However, it just doesn’t work due to The Bitter Lesson. Basically everything in classical NLP (machine translation, summarization, chatbots) lost to simple scaling. ChatGPT is a prime example—it used nothing from chatbot research and certainly wasn’t the intended end goal of OpenAI’s 2022 research program, but was a huge hit because someone (John Schulman et al) figured out the right way to package large language models as a product.
Method-driven research feels less stable because you’re constantly searching for problems and you have to be opportunistic. But I believe AI will allow method-driven research to dominate progress in most fields of science, one-by-one. The latest method (or “hammer”), as we’ve seen in AlphaEvolve, is ruthless search and optimization against a reward function (whether this requires RL or not is a separate discussion). Things that problem-driven researchers have been trying to solve for a long time like the kissing number problem will become nails hit by the hammer. Eventually the hammer will become bigger, stronger, and more general and will hit more and more nails.
So a very important meta-skill for the next decade will be knowing how to create the right environments to use The Hammer. Ironically, the problem-driven researchers, who by definition are experts in a specific problem, are well-positioned to create these environments. If, that is, they can put down their egos and pick up the hammer.
Confused about recent LLM RL results where models improve without any ground-truth signal? We were too. Until we looked at the reported numbers of the Pre-RL models and realized they were serverely underreported across papers. We compiled discrepancies in a blog below🧵👇
🚨 Paper Alert
🚨 Benchmark Alert
🚨 Dataset Alert
Sharing my last work during my tenure in Huawei, a robust benchmark and training dataset focusing on retrieval for long documents. Aside from page-level, we also introduce layout-level retrieval task to evaluate multimodal 📝
We show that :
1⃣ Visual retrievers consistently outperform text-based in page and layout-level
2⃣ Visual retrievers fine-tuned on our training set significantly outperform off-the-shelf models
3⃣ Token-level retrievers show marginal improvement over dense embedding models
Our training set is collected from seven DocVQA-related datasets with wide diverse domains. The benchmarks has 10 domains. We perform manual annotations for page-level labels and use external tool for layout-level labels
We only need ONE example for RLVR on LLMs to achieve significant improvement on math tasks!
📍RLVR with one training example can boost:
- Qwen2.5-Math-1.5B: 36.0% → 73.6%
- Qwen2.5-Math-7B: 51.0% → 79.2%
on MATH500.
📄 Paper: https://t.co/oROrmIHVwR
💻 Code: https://t.co/ao8ci5n66p
(1/n)
This "Aha moment" in the DeepSeek-R1 paper is huge:
Pure reinforcement learning (RL) enables an LLM to automatically learn to think and reflect.
This challenges the prior belief that replicating OpenAI's o1 reasoning models requires extensive CoT data. It turns out you just need to give it the right incentives.
We are so back in the AlphaGo excitement era: by playing countless Go games and maximizing the reward function (winning the game) using pure RL, AlphaGo beat the best human players.
Now we are entering the LLM RL era.
2025 could be the year of RL.