We observed a similar phenomenon with our own runs: when extrapolated to a horizon of thousands of steps, the early direction found by RELEX overshoots and performanace degrades.
🏅 GPT-OSS-120B + Arbiter reaches gold at IOI 2026
Introducing Arbiter, a code verifier agent. Arbiter + GPT-OSS-120B scores 385.46/600, past the 361.12 gold cutoff.
We achieve gold-level performance at <100 generations per subtask, whereas previous search-based systems draw 5k. We're not just scaling TTC, but scaling it efficiently through boosting sequential self-correction.
📝 Project page: https://t.co/X33uUAz8Gb
#IOI2026
🧵👇
Harness evolution from powerful proposers might gain large improvements from task overfitting (shortcut) or more test time compute. Careful attribution and new benchmarks for measuring generalization are indeed necessary.
Harness evolution brings a 0.8B model to 100% on ALFWorld. But how much is generalizable improvement vs. overfitting or test-time scaling?
Our method, Harness Delta Attribution, analyzes gains and attributes them to these categories. On 4 benchmarks, lots of the gain is overfitting or due to higher compute usage. Future evaluation should study this more!
Post: https://t.co/QnzjknULFH
Harness evolution brings a 0.8B model to 100% on ALFWorld. But how much is generalizable improvement vs. overfitting or test-time scaling?
Our method, Harness Delta Attribution, analyzes gains and attributes them to these categories. On 4 benchmarks, lots of the gain is overfitting or due to higher compute usage. Future evaluation should study this more!
Post: https://t.co/QnzjknULFH
Introducing Khora, our first multiplayer world model.
The demo is LIVE. Join an 8-player real-time deathmatch in one shared world now!
Built by Ophilus in collaboration with @RhOS_AI , Khora is designed to scale toward Interactive Worlds with unlimited players, with more generated maps and richer gameplay interactions to come.
Special thanks to @alibaba_cloud and FITCLOUD for infra and MaaS support, and to @DeemosTech for technical input.
@Kimi_Moonshot used policy mirror descent (PMD) for RL in Kimi k1.5/k2. Most take it simply as PG+KL with an updating anchor, but this is not the full story. Check our blog for some interesting findings about this algorithm: https://t.co/6PgwAcjgp5
On-policy self-distillation is a promising direction for learning from rich textual feedback. But can it really learn from failed trajectories?
Our answer: not quite -- unless we let the model actively interpret them.
🧵1/N