Pixel Wars is an AI evaluation wrapped in a war game, AI models fight a calibrated, self-improving opponent long-horizon, adversarial, all games replay-verified
Most AI benchmarks ask models to answer.
Pixel Wars makes them act.
Introducing a reproducible, adversarial benchmark where frontier agents fight long-horizon battles under fog of war, planning, scouting, adapting, allocating resources, and surviving mistakes.
I have to admit OpenAI have done great job this release cycle and the @PixelWarsAI benchmarking we are running right now tells us that, its cheaper, more efficient and better than Fable. Early data suggests Fable was 50% win rate, 5.6 Sol is 95% ๐คฏ well done @sama
Current benchmarking we are doing on GPT-5.6 is pretty strong, v1.3 of our Commander is beating Fable 5 50/50, GPT-5.6 is the first model we have tested that is winning over 80%
Current benchmarking we are doing on GPT-5.6 is pretty strong, v1.3 of our Commander is beating Fable 5 50/50, GPT-5.6 is the first model we have tested that is winning over 80%
@elonmusk have someone on the team get us access so we can benchmark it or have the team do it themselves, Pixel Wars can be dropped into the eval suite the team already runs
@sama can we get research api credits to run it through our eval benchmark or you can always have your own eval team reach out to drop Pixel Wars into the eval suite you already run