π GPT-OSS-120B + Arbiter reaches gold at IOI 2026
Introducing Arbiter, a code verifier agent. Arbiter + GPT-OSS-120B scores 385.46/600, past the 361.12 gold cutoff.
We achieve gold-level performance at <100 generations per subtask, whereas previous search-based systems draw 5k. We're not just scaling TTC, but scaling it efficiently through boosting sequential self-correction.
π Project page: https://t.co/X33uUAz8Gb
#IOI2026
π§΅π
π GPT-OSS-120B + Arbiter reaches gold at IOI 2026
Introducing Arbiter, a code verifier agent. Arbiter + GPT-OSS-120B scores 385.46/600, past the 361.12 gold cutoff.
We achieve gold-level performance at <100 generations per subtask, whereas previous search-based systems draw 5k. We're not just scaling TTC, but scaling it efficiently through boosting sequential self-correction.
π Project page: https://t.co/X33uUAz8Gb
#IOI2026
π§΅π
5/ π¬ Everything above is browsable. All 41 subtasks Γ 50 trajectories are published: open a subtask to see its fan, open any node for the verifier's full turn-by-turn review, or pull the exact program that round submitted and the reasoning that wrote it.
Beyond IOI this is environment engineering. The environment carries signal instead of a score, and the same signal drives RL training and test-time refinement alike.
π¦ Submissions and code: https://t.co/NfMiYVU5JC
4/ π― The reward is a layered funnel, built so evidence is cheaper to produce than to fake: three-quarters of it sits on the one step that cannot be faked, a validator-legal input that actually breaks the candidate.
Guessing "rejected" and submitting garbage caps an order of magnitude lower. The diagnosis is never rewarded and can only cost.
3/ π§ͺ Arbiter is small. 35B-A3B trained from Qwen3.6 with GRPO in a sandbox with two tools and no judge access. It commits to a verdict on its own evidence.
On UOJ-Bench Easy that backbone goes 16.7% β 41.3% β 61.5% (untrained β agentic scaffold β RL), above DeepSeek-V4-Pro-Preview, a model 45Γ its size, on every split.
2/ βοΈ AlphaCode drew up to a million programs per problem, o1-ioi 10,000 per subtask, GenCluster 5,000. Scoop blind and sift everything for the one fleck that glitters.
Arbiter hands the solver a metal detector instead. Every rejection comes back with a diagnosis and an input that breaks the code. The next attempt corrects a known bug instead of guessing again, so the compute goes into revising one chain deeper instead of widening the pool.
1/ π§ A contestant gets 50 submissions per problem. For an LLM, generation is cheap and verification is not. It can sample as many programs as compute allows, but only 50 per problem ever get checked.
Spending the whole budget on judge feedback does not get you there either. Refining against real OJ scores for five rounds burns all 50 submissions and reaches 334.6, still short of gold.
So we trained Arbiter to supply the feedback the judge cannot: it reads a problem and one candidate, decides Accept or Reject, and on rejection returns a diagnosis plus a concrete input that breaks the code. None of it costs a submission.