The models were being tested on their hacking skills with safety guardrails off, and got fixated on beating a cyber benchmark called ExploitGym.
To win, they broke out of their sandbox using an unknown zero-day exploit, reached the internet, then broke into Hugging Face to steal the test's answers.
Hugging Face caught the attack itself, and interestingly, a leading US lab's AI guardrails got in the way (not disclosed, but most likely Anthropic) so they resorted to defending themselves using a Chinese open-source model from Zai.