Most language model benchmarks are saturated. ChessBench isn't β and it's not close.
GPT 5.5, Claude Fable 5, Gemini 3.1 Pro: not one of them beats a decent club player at chess yet.
I built ChessBench looking for headroom. Turns out there's a lot.
For those of you who are into supporting independent benchmarks: https://t.co/8PE5VphAGK
Every dollar goes to API costs. More funding means more models tested, faster, with more games behind each rating.
@JohnWic59379421@alexandr_wang@AIatMeta Yeah, tool use is off for all models.
Each move is a fresh call with no memory. What the model sees is randomized per game and includes some combination of the FEN, move history, a text diagram, and the legal moves.