GPU parallelized envs have accelerated RL, but most implementations exhibit critical instability when running on-policy RL with short rollouts. We present Staggered Environment Resets. A few lines of code are all you need!
Presenting today, 4:30PM poster 310 #NeurIPS2025
🧵(1/8)
next-token prediction may be a sufficient objective (no RL postttraining) for general intelligence if neural networks didn’t bias for locality over globality (https://t.co/WZHeS0WzaQ). at some threshold a better compressor is more about understanding the underlying reality that led to the next token. poor compression in today’s LLMs may be because pretraining is currently done on i.i.d shuffled samples of data (multiple passes over), rather than some natural MDL-based curriculum (similar to how humans learn).
people are sleeping on the “multi-agent” aspect of the IMO/IOI results by @polynoamial@alexwei_. my guess is that the same model proposes tasks, generates candidate solutions, and verifies them. At a given iteration, if you can conceptualize and verify a problem of hardness x + epsilon and generate valid solutions of hardness x, you can bridge this epsilon gap iteratively without external verifier information.
@MechanizeWork my impression is that the larger risk of RL'ing models on low-quality envs is reward hacking, rather than compute wastage. changes in models' weights from RL-ing on poorly specified reward propagate far deeper than SFT-ig on low-quality human data.
@jchencxh is this provably private tho? even with a sparse/vague reward from evals from a customers employees, couldn’t you construct some setting in which RL can introduce sensitive information into the weights? tbf i think this would be very rare / afaik hasn’t been studied enough
@jchencxh is (2) feasible with privacy concerns? do large customers want these generalist models being RL’ed on their internal workflows? my impression was the rise of the RL FDE business is partially attributed to this
I’ve been bullish on learned tool usage for a while, especially for reasoning on images. Excited to see this in the latest OpenAI reasoning models!
I’m presenting a similar work on learning tool usage through evolutionary search for multimodal reasoning at #ICLR2025 in Singapore next week. Would love to chat with anyone on: diffusion models, symmetry learning / science of DL, RL, robotics, and more.