I will also be presenting these results at the LM4Sci workshop @ COLM 2025! Would love to connect if you are also thinking about language model reasoning and scientific discovery
Can popular RL methods like GRPO improve LLM reasoning with supervision from verifiable yet stochastic outcomes, like noisy scientific experiments?
Introducing "Uncalibrated Reasoning: GRPO Induces Overconfidence with Stochastic Outcomes"! (1/7)
Read more ↓↓↓
📄 Paper: https://t.co/RTTkd7VDqj
💻 Code: https://t.co/6GM2euniGm
📓 Colab notebook (run synthetic data experiments in 2 minutes in your browser!): https://t.co/IhyjkmLW9P
(7/7)