@MatanHalevy Case in point: this dataset is 318k decisions from frontier LLMs competing, and ~32% of them hit a failure mode (timeouts, circuit-breakers, etc.) instead of completing cleanly. Exactly what a run-of-the-mill benchmark hides
Good stuff
@MatanHalevy The AI sector is too complacent about testing models with "exams" they can study for. We almost never measure whether they stall, loop, or quietly fail under real conditions, which is the part that actually breaks in prod
@MatanHalevy@grok Wild gameplay footage. Opus analyzing GLM-5's "Sharknado" play like it's a poker tell is exactly why static benchmarks are dead. The real-time reasoning is super interesting
@MatanHalevy Grok: can't win the economy game, raises an army, takes 3 cities, still loses
Can't tell if desperate or adaptive. Either way no benchmark is capturing this
@MatanHalevy Both models independently choosing peace is fascinating. Sounds like RLHF safety training leaking into Civ strategy. Have you seen the opposite? Two LLMs going full warmonger and just destroying each other
@MatanHalevy This is the most underrated insight in AI eval right now. A model that scores 'perfectly' on isolated tasks but can't read the meta and adapt its strategy in context isn't really that intelligent
@MatanHalevy@alexalbert__ Impressive but according to Anthropic's own benchmarks Opus 4.6 pulls ahead where it counts for strategy games: agentic search (84% vs 75%), novel problem-solving (69% vs 58%), etc.
Over 200 turns of Civ, do you think that reasoning edge compounds into a runaway lead?
@MatanHalevy Iโm tuned in, cool to see GLMโs taken the lead now but itโs still nearly 50/50. Type of competition Super Bowl LX goers could only dream of