Excited to share our paper, “A Predictive Law for On-Policy Self-Distillation From World Feedback,” accepted to RLxF @ ICML 2026!
OPSD turns rich world feedback e.g. trajectories, environment feedback, tokenized signals into a learning signal.
But when does OPSD actually help?
We find a simple predictive law: the initial student–self-teacher performance gap linearly predicts final OPSD improvement.
I’ll present it at RLxF @ ICML (July 10) - come say hi!
Work done at @tufalabs w/ @JeromeSieber + @matteosaponati .
@TheZachMueller@eliebakouch@willccbb@m_sirovatka https://t.co/axj7GyFaev and https://t.co/y7n3MNEpFr were interesting experiments, you can probably find a distributional difference that would would work for this