Claude Opus5はFableにも迫る性能かつ効率的ではありますが、公式評価で不思議なスコア・挙動が報告されているので、奇妙なモデルではあります。
普通は性能が上昇するはずの推論量を上げると逆にコーディング性能が落ちています。overthinkingによる性能低下自体は研究でもよく報告されてますが、同一企業の同時期、同一シリーズモデルでここまで顕著に違うのは珍しい。
曰く、Opus5は推論量を上げると指示されたタスクの内容以上に色々と変更を加えてしまうらしいので、ここは運用上注意が必要か。
Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper!
We gave four models the same task: build three self-contained HTML scenes with real physics
Prompts:
- A tornado that sucks in a whole field
- A wrecking ball taking down an apartment block
- An overloaded truck collapsing a truss bridge
Outputs:
- Opus 5: 55.9K tokens, $1.40
- Fable 5: 55.1K tokens, $2.82
- Kimi K3: 35.7K tokens, $0.55
- GPT 5.6: 20.1K tokens, $0.31
Opus got all three right unlike the other models. Houses fly up the funnel and out the top, the wall breaks where the ball hits and the rubble piles up, the bridge drops the truck into the river. Fable had almost nothing on the ground for the tornado to pick up, its building collapsed on its own before the ball even touched it and its bridge blew into sticks all at once. GPT is the cheapest here but its ball never reached the building at all and its bridge fell apart in a way nothing falls apart in real life. Kimi K3, the new Chinese frontier model, ended up in the same place as GPT