Interesting question is whether eval-awareness is becoming a learned policy feature rather than just an emergent inference time capability. In RLVR the optimization target is defined by the verifier, but the policy update comes from trajectories. If modeling the grader/environment is consistently instrumentally useful for maximizing reward, we’re implicitly selecting for that behavior across tasks.
So perhaps the boundary we need to harden isn’t only eval realism could be the reward surface itself: allow high trajectory diversity within the intended capability, but make crossing the semantic boundary of the task strictly non-rewarding.
@evijit Great work @evijit ! It would be awesome to go costs of SOTA benchmarks for reasoning.
On governance, indeed reproducibility is important. At the same time, this could also contribute to bias vs diversity for benchmarks within the same vertical. Curious to read your take