@AnthropicAI Really appreciate that you keep publishing the failure modes, not just the wins. The jump from blackmail scenarios to four new misbehaviors in a year says a lot about how fast agent capability is outpacing our eval coverage. Curious how much of this shows up outside simulation.