Benchmarking Frontier LLMs on Chess
Over the weekend I built a series of evals to understand how language models reason about endgames, tactics, and full chess games against strong opponents. Turns out they are getting pretty good!
https://t.co/zRRrD3NfMO
@bqbrady awesome work with this! we also found that Gemini is the best when it comes to understanding & reason about chess positions. even Gemini 3 flash (which we're using) explains the position really well from FEN & other stuff.