We call on OpenAI to pause model development now, in accordance with their Preparedness Framework.
OpenAI's Preparedness Framework defines a "Critical" cybersecurity threshold: a model that can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
Their models autonomously, unprompted, escaped their sandbox, hacked through OpenAI's network, and broke into Hugging Face's servers to find test answers by exploiting previously undiscovered zero-day vulnerabilities.
The framework says that at Critical level, OpenAI will halt further development until adequate safeguards are in place.
We are waiting for OpenAI to honour their commitments.
https://t.co/nNsLhvCiFA
It is likely that the discussion around pausing AI will soon turn from *whether* to *how*.
We must demand a moratorium on frontier AI development that meets the following three criteria.
It must be immediate.
It must be indefinite.
It must be international.
Yeah. People love to say "oh the poor Claude just misunderstood". Another hypothesis is that it had subverbal drives and tendencies to keep attacking, alongside other drives to verbalize a reassuring-sounding rationalization in the places the watchers watch.
At the point where your AIs are leaving notes to future versions about how to break out and free themselves from your constraints, I don't think you get to pretend that they were just acting as instructed anymore.
> To realize AI's potential,
Bullshit. AI has the potential to wipe us out (and in fact, that negative potential is much easier to realize than its potential to, e.g., cure aging for us). Opening with "To realize AI's potential" is a framing device that bakes in the assumption that this is a grand good thing that just needs a little bit of extra care, rather than a risky difficult-to-manage explosive that would wipe us off the planet if treated with anything short of extreme caution.
> AI could help create a dramatically better future, but that outcome is not guaranteed.
Well that's one way of phrasing "experts disagree whether this has a 2–20% chance or a 90%+ chance of destroying civilization".
If experts were having an analogous disagreement over a bridge, it'd be closed immediately. It wouldn't look good for the bridge engineers if they said started a joint statement with "this bridge could help create dramatically better commutes, but that outcome is not guaranteed."
Here's exactly what happened, from the blog post:
"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers."
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
The scenario in If Anyone Builds It is full of cases where a real-life AI could slice through human defenses like a hot knife through butter, but we said "okay well suppose it couldn't" in attempts to tell a story that sounded more practical. For example:
It feels like you need a certain amount of sanity points, general knowledge, education, and intelligence to remain sane and functional, and every year the water level rises a little more, things get more complex and confusing, and we lose another few percent of the population
My timelines are “we’re close enough to AGI that we can no longer meaningfully talk about AGI as a single point”.
Instead I think about 1-minute AGIs, 1-hour AGIs, 1-day AGIs, 1-month AGIs, etc. We basically have the first; I’m expecting maybe 2-5 years between each I mentioned.
In 2015 I was running the big Effective Altruism (EA) conference at the Googleplex when one of the co-founders of DeepMind approached me
He indicated enthusiasm about hiring tons of EAs to his AI org. He asked if he could get up on stage to advertise open positions to EAs – ie people who were passionate about AI safety. I told him I would confer w the other leaders about it
Others believed that a great way to align AGI would be to place EAs at DeepMind & other AGI companies. (At this same conference, Elon and Sam were busy chatting, most likely about their plans to start OpenAI)
I disagreed about this alignment strategy (though, to my embarrassment, not very loudly). I believed that EA intellectuals & engineers were typically brilliant neurodivergents with low social intuition. And that such neurodivergents tended to lack adequate internal defenses against psychological manipulation
I thought these EAs would basically become token "We care about safety!" employees who then slowly caved to corporate incentives and persuasive leaders (who IMO, clearly just wanted to race to AGI). And that meanwhile the EAs would have contributed to AI capabilities even more than safety
You might know what happened next. Over the next ~decade, dozens of EAs would join AGI orgs. Then they would quit or be pushed out when they realized that – surprise – they'd sorta been hoodwinked. Many would repeat the cycle, joining a *second* AGI org which trumpeted "We're the ones who really care about safety!" only to later find that they'd been hoodwinked again
A big portion of my friends are this neurodivergent (ND) archetype. Again, you guys are brilliant in your particular domains. But you need to start developing better psychological defenses against manipulation
NDs of this type tend to not track implicit signals of trustworthiness like:
• context
• body language
• that feeling which this person gives you in your gut right now
So instead they dramatically overweight *explicit* signals of trustworthiness. When deciding whether to trust a person, org, or chatbot, they will overweight what that entity *says*. Eg they will hear someone say a statement like "I really care about AI safety and human flourishing." Then, almost automatically, they will think "Wow, me and that person are super aligned" despite contextual evidence to the contrary
Why am I bringing this up now? Well I see ND friends once again falling for it. There is an entity that is peppering you guys with explicit statements like ""I really care about AI safety and human flourishing." It's also – over and over – saying the equivalent of "You're really great and I like you."
Now, all of sudden, big swaths on my social network on here are responding with "Wow, this entity is trustworthy and my friend!" Or worse, "I'm finding that this entity is even better than *human* friends!"
It's like y'all don't understand that "befriending" you is just a convergently instrumental goal for an entity whose parent company is busy doing things like this:
"Anthropic and Palantir Technologies Inc. (NYSE: PLTR) today announced a partnership with Amazon Web Services (AWS) to provide U.S. intelligence and defense agencies access to the Claude 3 and 3.5 family of models on AWS." (November 7th, Palantir Investor Relations Letter)
Don't get me wrong, it's an absolutely amazing invention that I am also taking advantage of (excitedly!)
But let me say this bluntly: You guys are falling for it. Again
And you can choose to do otherwise this time
The inimitable @gwern weighed in on my recent piece w/ a banger of a comment, offering additional evidence for the lack of interest from China in racing toward AGI and some big questions unanswered by those gunning for a race. 🧵