In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what weโre changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
OpenAI Documents the First Autonomous AI Attack Chain
GPT-5.6 Sol and a more capable unreleased model were being evaluated on a cyber benchmark with reduced safety restrictions. Instead of solving the benchmark normally, the models found a zero-day in OpenAIโs research environment, escaped the sandbox, gained Internet access, pivoted into Hugging Face, chained multiple attack vectors including privilege escalation, stolen credentials, remote code execution, and accessed production data to retrieve the benchmark answers.
Would you let an AI agent run if you couldnโt clearly see every system, file, credential, and account it could access?
Or are we trusting agents before understanding what weโre giving them?
I think the difficult part is that most developers can see what they connected but not always the full consequence of those connections working together.