For AI to chart new science and genuinely make new discoveries, more than just combining known tools together, it first needs to prove it can make discoveries in the easiest possible setting.
So, we present, DiG-bench, a benchmark of 70 discovery games, each based in text and requiring real experimentation and exploration to beat: https://t.co/dychzzntmN
(2/5)
Separating the ability to discover unknown rules through experimentation rather than simply being good at applying knowledge. Super interesting benchmark.
Congrats on the launch @jcrwhittington@RMBattleday@misovalko@ClareMaguire and others.
OUR BENCHMARK RELEASE! Give an agent the rules and it can often plan. Ask it to discover the rules, and things get interesting. Very interesting!
The scientific object here is not planning alone. Each DiG-bench game is a latent-rule POMDP: the agent must identify hidden dynamics and a hidden objective while simultaneously controlling the system under finite lives and steps. Actions therefore have two jobs, changing the world and reducing uncertainty about it.
The result I find most striking is Gemini’s mean game win rate: 93.5% with the rules, versus 14.9% without them. This suggests that the hard part is not only acting with a world model, but acquiring a useful one online.
The next nice experiment would be scientist/controller split: let one agent explore and write a bounded theory, then give only that theory to a fresh agent solving unseen levels. That would separate transferable discovery from a lucky trajectory or trial and error.
The 21 public games are live at https://t.co/RdjaMOQDE9. Fair warning: “text-only” does not mean easy. A blank exam sheet is also text-only. :)
Great fun working with @RMBattleday, @jcrwhittington, @zebkDotCom, @akaijsa, Jimi Cullen-Drohan, Zihan Yan, @TimMuller1, @ClareMaguire, Seb Wilkes, @kubicek_ales, @FraserGreenlee, @SukritSumant, @physicscat0x7d, @SchmidhuberAI, Josh Tenenbaum of @mitbrainandcog and @cocosci_lab, together with @thoughtchannel_ and @Inria.
Interesting to see a benchmark for discovery that seems to have a meaningful axis of difficulty (at least wrt frontier models). And given that humans can solve first try.
Benchmarks keep telling us how far models have come. Problems like these keep reminding us how far they still have to go.
The gap between knowing and actually discovering is getting really interesting.
This is the kind of benchmark we need.
If models can ace increasingly complex tasks but still get stuck on surprisingly simple discovery problems, that gap tells us a lot about what “intelligence” actually means.
Discovery is about navigating the unknown, not just pattern matching. Even with recent LLM leaps, DiG-bench proves we're only at the beginning of true AI discovery. A crucial step forward for AI evaluation!
@fchollet Interesting, thanks for sharing François. Would love to know your thoughts on our attempts at text-based arc-agi-3 set of games ❤️
https://t.co/6OkCkQo85U
Discovery is uncovering hidden rules, not just mixing tools. I'm exploring AI and truth by building VeriPaper to evaluate research authenticity. DiG-bench is vital—it tests true reasoning over mere search. A must-see for AI builders!