A bit of a history of IMO-Bench and our IMO efforts:
a. We started building IMO-Bench around early 2024, which was the precursor of ProofBench (basic).
b. IMO-Bench was first mentioned in the Gemini 1.5 paper around May 2024. At that time, Gemini Math-specialized 1.5 Pro scored only 25% (whereas the performance on Hendrycks’s MATH was 81%, a big breakthrough!).
c. Fast forward to April 2025, Gemini 2.5 Pro scored 55% on ProofBench (basic) and 35% on ProofBench (advanced). Nobody talked about the MATH benchmark anymore (it has served its purpose well!).
d. Also in April 2025, our paper was accepted to ACL 2025. We were asking leadership to share it publicly, first to arXiv, but we were asked to wait until IMO 2025, so we postponed ACL. At the time, it felt like a major setback to me because it was uncertain at the time if we were going to surpass ourselves at IMO 2025 (we got Silver at IMO 2024). We kept marching on.
e. At IMO 2025, our generalist model (advanced Gemini Deep Think), scored 89.0% on ProofBench (basic) and 66% on ProofBench (advanced). And the rest was history 🙂
In case people missed it, this project page has all the infos about IMO-Bench https://t.co/jB67vqApu3.
DeepMind has ambitious plans to make massive generative models that simulate the world. I'm hiring for a new team with this mission. Come build with us!
https://t.co/pqvALtAvLs https://t.co/vtwgeXl9Dl
@_ianwalker@O2academybrix@O2@EatYourOwnEars@FourTet It was just poor sound and far too loud. London venues are way beyond that. It’s not Fabric Room 1 but Brixton Academy usually has a great system. Shows the value of decent sound engineers!
@EatYourOwnEars@FourTet @caribou The sound was terrible. I have tinnitus now because of it. Loads left before FourTet came on because it was so bad. Why did you not have anyone at the sound desk? Feel bad for the artists who rely on the venue & promoters to sort this stuff out #tinnitus#basics#seewhatyoucando