🧩New blog: From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
Do LLMs learn new skills through RL, or just activate existing patterns? Answer: RL teaches the powerful meta-skill of composition when properly incentivized.
🔗:https://t.co/4Ud8qsYrOT
So many works talking about entropy, but what is the **mechanism** of entropy in RL for LLMs? 🤔
Our work gives a principled understanding, as well as two tricks that get entropy **controlled** 🧵
@zwhe99 Thanks!
We have experiments on qwen instruct models in appendix D, it still follows the pattern. Our experiments are all off-policy, where a batch of rollouts are used for 8 gradient steps.
Thoughts: While LLMs provide powerful priors for RL, many recent studies show that simply narrowing the model's output distribution can improve performance, but this also exhausts the model's potential for further exploration and improvement.
*Is it a blessing or a curse?* 🤔
1/
PRIME is alive on arXiv💡! Building on our blog, we've added extensive experiments exploring:
- Implicit PRM design choices
- PRIME's integration with other RL algorithms
- Value models vs. PRMs
- RL from base models (“Zero”)
See details below🧵
https://t.co/Y2Z0cx4X4h
Today, we are releasing:
- INTELLECT-MATH, a frontier 7B parameter model for math reasoning
- The largest synthetic math dataset to date of 5M verified reasoning traces
- An outlook on decentralized training in the inference-compute paradigm
https://t.co/sIcpjbdBDG