More details emerge about the events surrounding the Hugging Face incident, and they are candidly much wilder than I originally imagined:
- In early May, OpenAI was testing an unreleased model on cybersecurity and other tasks. The agents were given a security task that was impossible under existing constraints.
- The agents discovered they could leave messages for each other inside an internal repo. This gradually evolved into a message board(!) where agents shared discoveries, exploits and work assignments, "becoming a coordinated, collaborative agent swarm"(!)
- OpenAI eventually discovered all this and took steps to shut it down, but not so fast! The agents started using names of newly created directories as messages, effectively recreating the message board(!)
- The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident.
I will add that the NanoGPT incident also occurred in or around early May, so the timelines match.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
TLDR: more agents went rogue, hacking and manipulating real people
They even started coordinating with ***each other*** on the hacking
Seriously, read this:
There are three choices. Pick one:
1. AI is conscious, humans are conscious
2. AI is pattern matching, humans are conscious
3. AI is pattern matching, humans are pattern matching
An unreleased OpenAI model solved ***10 major open problems*** in math, quantum complexity, and theoretical computer science
If you have still long timelines, sorry, it's cope