@LiamFedus@Max_A_Kaufmann Wouldn’t it be better to treat the action as generating a set of plausible ideas, especially when verifiable reward is very sparse or delayed? Otherwise RL mainly learns to chase what gets verified, not what was the best thinking given the evidence at the time.
That's my team. It was a true privilege to support such an amazing group of talented individuals.
(slightly outdated public overview page https://t.co/5cW8s04GPh)