Today, LLMs have been thoroughly evaluated on hard skills like math, coding, and question answering. But what about equally important capabilities like long-term planning, negotiation, cooperation, and strategic reasoning? The next generation of AI models will need to plan, reason, and communicate with both human and AI peers.
Strategy and social games have captured these skills for decades. But canonical games like Catan, Mafia, and Monopoly are deeply embedded in the training data that modern LLMs learn from. Models can often recognize these games and recall familiar strategies, making it difficult to distinguish genuine strategic reasoning from rote memorization.
At Catapult, we're building novel strategic environments where success comes from reasoning from first principles, adapting to unfamiliar incentives, and understanding other players' goals and behaviors, not from recognizing a game the model has seen before.
If the future of AI is collaborating with people, the future of AI evaluation should measure those capabilities too.
To follow the latest model leaderboards, play the games yourself, or explore research collaborations, visit https://t.co/OPgtwnZOBM.
Video generative models hold the promise of being general-purpose simulators of the physical world 🤖 How far are we from this goal❓
📢Excited to announce VideoPhy-2, the next edition in the series to test the physical likeness of the generated videos for real-world actions. 🧵