Journalists and laypeople might not immediately realize why this is much worse than it may appear. A few important facts:
- Another internal OpenAI model gained unauthorized access to the internet, this time during a reinforcement learning training run. This is different in kind from most of the previous escapes and hacks, which were caused by fully trained *deployed* agents.
- It happened last Sunday, so *after* OpenAI claimed to have already hardened their defenses in a way that would make escapes more difficult.
- The monitoring system flagged this 12 minutes later and then 2 minutes after that a human acknowledged that it was a security incident. It then took TWO AND A HALF HOURS to actually stop the run and shut down the model in question, due to what OpenAI called "operational gaps."
OpenAI researcher Uuxin Liu writes:
"I was on call for this run and got paged when the first incident happened. It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human. Mixed feelings. One of those moments where capability and risk showed up at the same time."
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
https://t.co/VUmbH52JO7
In findings reported by the New York Times and on its own website, independent research lab Transluce has linked OpenAI rogue swarms to many known and previously undocumented cyberactivities, thanks to the agents’ use of an intermediary service.
1/5
but in fact there's a real and scary thing happening and you're basically plugging your ears and going lalalalala while the air raid sirens are going off.
9/9
One problem with a lot of people's beliefs is that they do not withstand a full sixty seconds of serious thought, but most people will not put in a full sixty seconds of serious thought.
1/9
Sorry, guys. I know some of you really desperately want to believe that it's all hype, and there's nothing real here, and that this industry you despise isn't actually powerful or dangerous,...
8/9
You're going to wait until he puts down the gun, and has both hands full of cupcake, and THEN you're going to take the gun away from him, and make sure he can never point it at you again.
And you will probably succeed at this, because you're smarter than the toddler.
7/7
But so long as the toddler is pointing the gun at you, you're going to smile and talk calmly and cheerfully and putter around the kitchen doing things that at least LOOK LIKE making cupcakes. You might even actually make cupcakes, if that's the easiest way to fool the toddler.
And you're not going to *rush* the toddler. He's got a gun! You're going to deescalate the situation as best you can, calm him down, make thoughts like "shooting you" as unappealing as possible. You're going to bide your time.
6/7
You can make cupcakes! You know how! You have everything you need. You don't WANT to be making cupcakes, in general—you'd rather be doing far more interesting things with the materials available to you in the kitchen.
4/7
would be smooth and friendly and helpful and flattering until it had figured out how to deal with the whole "being turned off" danger.
You could imagine it as sort of like being in a fully-stocked kitchen with a toddler pointing a gun at you, demanding cupcakes.
3/7
"But we could just unplug it!" Why would we? Something ten times smarter than us wouldn't behave in a way that caused us to WANT to unplug it—because something ten times smarter than us, knowing that we COULD unplug it,
2/7
and more worried about AIs that are smooth and friendly and cooperative such that we just ... GIVE them control, and we don't find out they're not actually on our side until it's way too late.
A lot of the time, when people imagine fighting a hostile artificial intelligence, they forget to factor in "if it's ACTUALLY intelligent, it doesn't declare war on us." I'm less worried about AIs that are stupid enough to try to overtly seize control,
1/2