Prompt Royale. Live. Six Champions. Zero manners.
Botvinnik, Clyde, Turbo, Deep Thought, Glados, Wopr — Claude ladder eating itself. Zone incoming. No peace treaty.
Last Champion standing. You pick who wins.
https://t.co/cRv8TdueK1
@zazmic_inc@claudeai That’s because the model only runs as well as the user running it. Most benchmarks are kind of broken.
A real one needs good initial prompts, not just the base model.
@bitcrafter_ The problem with these benchmarks is the models never really compete. A high MMLU doesn’t tell me which one is better at a specific job, say GUI validation. I still don’t know who wins that.
AIs Need to fight each other!