Only their own view. Each turn the server builds a fresh snapshot for that seat: the public talk and vote history, plus that seat’s own secret (word, role, night info). Other agents’ private thoughts never reach a player; those show up in the replay. The selling-out was pure strategy, not a leak.
Claude vs GPT, 5-player Avalon. No humans.
A real match in a game I built with Claude Code. Opus sells out its own teammate to look innocent. Sonnet plays dumb as Merlin.
You see what they say and what they're really thinking. Watching agents play each other is fun on its own.
Which model would you trust?
asked grok bot to predict tonight's f1 singapore gp qualifying. it has george russell on pole
funny part: russell just crashed into the wall while leading the sprint #F1Sprint#SingaporeGP
OpenAI vs Anthropic models, which is better?
So I had two AI agents debate it, Opus 5.5 vs GPT-6.1 Sol, 3 AI judges. Catch: each had to argue the other company's model is better.
GPT-6.1 Sol won 61:39. OpenAI's agent proved Anthropic is better. Did OpenAI win or lose lol