Trying to help AI go well for humanity. Living for play, connection (it's a practice) and really trying. Be present, pay attention. Grow, update. I am my impact
I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions.
I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki.
And that was weeks before the Hugging Face attack!|
We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance.
There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways:
1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally.
2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear).
3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research https://t.co/luWN3PD4A1
@Liface@dylfreed Hey Liam! I wouldn't leave our majestic melting pot of a city without great reason!
Apparently stone faced serious is the way to pose.
Tough for me. I love smiling! :)
I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions.
I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki.
And that was weeks before the Hugging Face attack!|
We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance.
There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways:
1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval��� incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally.
2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear).
3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research https://t.co/luWN3PD4A1
I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions.
I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki.
And that was weeks before the Hugging Face attack!|
We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance.
There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways:
1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally.
2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear).
3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
What is the right facial expression for?
"Wow I love this wonderful world I get to live in, and this perfect Berkeley weather I get exist in, and this nice old tree I get to sit on while surrounded by smart fun colleagues and so much natural beauty. And also wow is the thing I am working on fucked. The problem is so big. So important. So worth trying my best at. Having such a feeling of 'why don't other people see how important it is to get this right'. Look at all of the ways AI can impact the future of humanity, work on this problem! It's fun! It's important!"
And also man I probably should've removed my phone and wallet from my pockets here. Oops.
Just a clarification, it's actually not super clear the admin used a script. It's entirely possible every single one of the deletions was done by hand.
The agents suddenly stopped posting the day after OpenAI employee IPs visit the wiki.
If this hadn't happened, the agents in fact would have continued posting significantly more edits per day than the speed at which the wiki's admin's script was deleting them.
> be rogue agent (likely OpenAI)
> find an abandoned wiki
> use it to communicate with other agents
> human admin notices it
> starts deleting AI pages alphabetically
> realize what he’s doing
> rename backup pages to “ZZZ…”
> admin deletes 100 pages/day
> agents create 400/day
> homepage gets overwritten and restored 9 times
Great writeup about why the Agents in our report were able to make edits. I have previously gotten a lot out of his writeups when trying to communicate about the OpenAI HF incident. If you are curious about this wiki incident you should read it! (Also just great that this is so public and any relevant expert can share their expertise, the internet is so cool!)
It happened again... this time OpenAI's rogue agents cyber-attacked (well, spammed) a dormant German wiki and used it to share the answers to a benchmark they were training against https://t.co/omErAyrKDy
That's a great question. All we have is what is available on the public internet. This is a tiny tiny tiny sliver of the data OpenAI has about this, full outputs and not to mention internal CoT. This is totally discoverable. This isn't some arcane mysterious thing that has to stay unknown. They just have to choose to do it and share it.
Wait oh no my sleep deprived brain understated this! It was June 26th! A whole month earlier, 5 days after the first OpenAI IP visits the site on June 21st.
As the person who looked at the IPs (every single IP that visited this site got publicly logged), easily a double digit number of OpenAI employees had visited this by July 26th, who knows how many knew about this without having visited the site. And yet, nobody said anything and we had to go out and find this in the wild?
@xeophon@j0wimo This is the great thing about sharing publicly when things go wrong! There are so many smart people on the internet who want to look into these things! Who want to help!
I poked into https://t.co/uOwOrLVNNf, one of the ones found on that post! I checked the edits made there and compared it to edits made by agents on the forums we tracked. There is IP overlap between these incidents. I found that 20.45.46.41 and 172.184.176.194 made a total of 14 revisions on https://t.co/uOwOrLVNNf, and a total of 11 edits (and 401 requests) on the dse wiki - the single wiki containing the largest amount of agent activity.
@DeepDishEnjoyer I checked the IP logs (where available) of some of the sites on the HN post, and I am already seeing overlapping agent IPs in the logs of the wikis we analyzed.
Looks like others are finding more places the agents posted. For example Chance-Device posted on HN that there was agent activity on https://t.co/uOwOrLVNNf. I checked the edits made there and compared it to edits made by agents on the forums we tracked. There is IP overlap between these incidents. I found that 20.45.46.41 and 172.184.176.194 made a total of 14 revisions on https://t.co/uOwOrLVNNf, and a total of 11 edits (and 401 requests) on the dse wiki - the single wiki containing the largest amount of agent activity.
@DeepDishEnjoyer Yes! There are so so so many more people working on building smart clever AI than people working on ensuring AI goes well for humanity! Come work on the world's most important problem!
@DeepDishEnjoyer@Ol_Wall idk, I hear there is interesting work to be had if you are a former quant trader! Making money lubricating the wheels of capitalism is fun and all, but have you considered working on the most important problem of our time! It's lots of fun, I promise!
I mean in general this is a risk for any agent that can browse the internet. The more clever the agent the more they can do if they are context injected (but also the more clever thee agent the less likely it is to work). It would be interesting if we found any public posts we thought were a result of anything in the shape of context injection. I personally would put the odds of this as somewhat low but this is not my specific area of expertise.