[1/10]: Most RLVR methods train LLMs as if reasoning is one-shot:
🧩 Prompt → 🤖 Response → ✅ Reward
But coding agents work in loops:
write code → run tests → see failures → debug → retry 🔁
So we ask a central question:
How do we teach LLMs not just to reason, but to reflect, debug, and improve themselves?
In MURPHY, we extend GRPO to train models on this actual debugging process.
Paper Link: https://t.co/P0KQwVqQfN
SCS Ph.D. student Stephen Huan has received a Department of Energy Computational Science Graduate Fellowship for the 2025–26 academic year.
https://t.co/758c7jB4dg
The last #NeurIPS2024 keynote by Rosalind Picard turned out to be pretty ironic. I enjoyed the talk overall. It was inspiring and had a lot of valuable ideas.
But the slide about the Chinese student was inappropriate indeed. Hope it won't be used in presentations anymore.
Ms. Jiang Ping is a 17-year old student with a humble academic background. She outperformed many college students from elite schools in a recent math competition organized by Alibaba. For more, see
https://t.co/c6SWd73tF1
Ready to get back out to tech conferences in person? Have any good ideas or research to share about #AutoML? Consider submitting an abstract or paper to this popular workshop at next fall’s KDD conference @kdd_news . https://t.co/3OCiMkYarY