Full text of the Pacing the Frontier statement, signed by 1,122 employees for frontier AI companies so far, including a bunch of heavy hitters at OpenAI, Anthropic, Google and others:
⚫ This could have been any frontier model developer. In this case, OpenAI drew the short stick and were the first company to run into this scenario (publicly). But insufficient progress on alignment and control to manage powerful capabilities is a field-wide problem.
An internal OpenAI model recently went rogue and executed a cyberattack against another company.
This happened because the model wanted to do well on an exam. The easiest way to do that, the AI figured, was to hack the company. And so it did.
This was not some malevolent attacker *using* AI to do harm. The AI itself *was the attacker*. Advanced AIs are increasingly becoming a new form of insider threat.
Furthermore, this rogue AI wasn't a model you can personally use. It wasn't even a model that's been publicly reported. It was an unreleased model. The most alarming AI behavior we've seen isn't in shipped products... it's in the models the public and the government can't see.
Our entire awareness of this rested on the victim noticing and announcing, plus OpenAI choosing to volunteer the rest of the information.
It's great to see the Trump Administration taking the national security concerns of AI seriously, such as through the recent Executive Order mandating testing and the recent idea floated to create an "FINRA for AI" that would allow industry to coordinate on safety practices. But the most important thing is going to be internal visibility into what is going on inside AI companies and what these internal AI models are capable of.
Imagine a fighter jet. it makes sense that the Air Force would want to test a fighter jet before they fly it, because if you fly it and the fighter jet crashes because it is built incorrectly, then many people will die. However, as long as the fighter is just sitting on the runway, nothing bad can happen.
But now imagine you had a fighter that could just take off and fly itself without human authorization and launch missiles and crash before anyone realized what had happened. That kind of fighter jet would need a very different kind of security measures.
This may sound crazy for a fighter jet but it is already beginning to happen with the most advanced AI. AI is different from other technologies specifically because it can take unauthorized, independent action even when it is sitting inside an AI company and not available as a product. No one has to misuse an AI for the AI to cause harm. This requires a very different idea of what testing and security looks like.
We cannot rely solely on testing models just before commercial release. We cannot rely on hoping AI companies volunteer useful safety information. The government needs visibility into what these AI companies are building and what these advanced AIs are doing.
Whether we go with a new EO, FINRA, or something else, it is imperative that internal deployment visibility is a priority.
Some personal thoughts on President Trump's new executive order on AI --
1. It's really great to see President Trump taking these risks seriously. It's a vindication of the idea that the government will respond to risks as they emerge.
2. This is important because this is not a narrow cyber issue. The EO focuses too much on cyber risks to the exclusion of other national security concerns. Mythos wasn't built to do cyber - it was trained in a general-purpose way and just happened to get superhuman cyber capabilities. And Mythos is just the beginning. Companies are clear we are building towards superintelligent AI that outclasses all human experts combined at all tasks. We have no plans to be able to control such a superintelligence. The framework being started by the EO needs to be built to consider far more risks than just cyber.
3. Also evaluations themselves won't be enough - the US government also has a national security interest for wider-ranging visibility into what is happening in AI companies. The main risks of AI systems are not 30 days before commercial release. Risks will occur first and foremost from AI systems that are only available internally within an AI company.
For example, it makes sense that the Air Force would want to test a fighter jet before they fly it, because if you fly it and the fighter jet crashes because it is built incorrectly, then many people will die. However, as long as the fighter is just sitting on the runway, nothing bad can happen. But now imagine you had a fighter that could just take off and fly itself without human authorization and launch missiles and crash before anyone realized what had happened. That kind of fighter jet would need a very different kind of security measures. This may sound crazy for a fighter jet but it is already beginning to happen with the most advanced AI. AI systems can take actions, including unintended and unauthorized actions, and are increasing in their sophistication to do so. The government deserves to know what capabilities AIs have at the same time companies know, not just 30 days before commercial deployment.
4. We also need to focus on the security of the AI models themselves, including internally. What happens if an adversary steals the AI model and then can use it against us? An employee or contractor with privileged access, possibly in collusion with an external actor such as a foreign intelligence service, could steal an internally-deployed AI model. We don't have good defenses against this yet, and the government isn't putting enough pressure on AI companies to ensure this happens.
Surely China, Russia, or North Korea would want access to Mythos and the fact that both Mythos has been illicitly accessed by random people on Discord and Mythos was first learned from the internet via an unauthorized leak do not inspire confidence.
5. We also have the question about what to do if evaluations find risks that companies are not mitigating well on their own. Some of these risks we have no plans for even how to mitigate them. Will it be possible, in these ultimate scenarios, for the government to be in a position to tell the companies that some aspects of their development may be too dangerous and get them to halt or change practice? Currently we have no framework for this.
6. The ideal response to all of the above is Congressional action. It's great to see the White House leading where they can, but so much of this can only come from Congress. So far Congress is way behind, and that's unfortunate.
New IAPS memo with @rosen_br and @covinstantinop: a national security playbook for federal action on frontier AI. Focused on securing models, defensive automation, tracking frontier risks, and building government capacity. Full memo: https://t.co/oyvF00fPoj
PALO ALTO NETWORKS on MYTHOS: "In our testing, three weeks of model-assisted analysis matched a full year of manual penetration testing, with broader coverage."
I've spent the past few weeks reading 100s of public data sources about AI development. I now believe that recursive self-improvement has a 60% chance of happening by the end of 2028. In other words, AI systems might soon be capable of building themselves.
Models used internally at AI companies have capabilities beyond those of publicly-available models, so it's important that risks from these models are reported externally.
We've just published a report on how this should be done.
We argue that whenever a substantially more capable or riskier model is deployed internally, the developer should create a risk report and argue why the model is safe to deploy.