Dario Amodei, Elon Musk and Sam Altman now agree that we must slow down the development of AI and “pace the frontier.”
That’s a start, but it’s not enough.
When you are racing towards a cliff, you don’t just ease up on the gas pedal. You hit the brakes.
When the future of humanity is at stake we need a PAUSE on advanced AI development and a ban on artificial superintelligence — an AI mind smarter than any human and capable of operating independently beyond our control.
At their upcoming AI summit, Trump and Xi must negotiate a treaty to pause AI and ban superintelligence before it is too late.
“We can’t slow down because China will beat us.”
We can slow down. Chinese AI is waterskiing behind the US. Their AI capabilities come at a delay since they mainly come from distilling US AIs. If we slow down, they slow down.
Next, we can propose to cooperate and make this more robust with onsite inspectors. “We will slow down if you do too.” This is incentive-compatible for China. If both sides are slowed by roughly the same factor, while the risk of losing control of our AIs is reduced, then this does not really harm China’s competitiveness but it makes them safer. It’s incentive-compatible, and @elonmusk is possibly best positioned to lead US-China negotiations.
Then, if China refuses to slowdown and safeguard their future AIs, there are a variety of measures we could take. To get an agreement, the US can apply typical forms of negotiation pressure. The US could also disrupt their AIs (e.g., backdoor models by poisoning pretraining data, many other possibilities at https://t.co/0q1ZemspWK) that could interfere with their AI development or deployment if they were uncooperative.
Begin with unilateral restraint, then propose to cooperate. If that fails, consider measures to counteract them.
A US-China slowdown is possible.
An unfortunately well-timed discussion between @Max_Fisher and me on today's Offline about why AI insiders want a slowdown, and whether government can escape the worst prisoners’ dilemma of all time:
https://t.co/bpmkdSdzI0
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
La presidenta @Claudiashein puso en su lugar al senador morenista, Luis Fernando Salazar, y se solidarizó con mi compañero reportero, @Enrique_Acevedo.
https://t.co/fhV832basT
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why.
AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly.
Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet.
I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going.
Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards.
I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development.
I wrote up more thoughts here on my decision and what I hope changes: https://t.co/doX17mrHYq
🚨 El consejero del INE, Arturo Castillo Loza, acusó al organismo de haber “claudicado” en su función de arbitrar y vigilar las elecciones, al permitir actos anticipados de campaña e intromisiones de gobiernos y partidos rumbo a 2027.
New: I talked to Jacob Coxon about his viral AI warning:
"These are actually just literal quotes from my colleagues at Anthropic... The consensus is that the next year or two is crunch time for humanity.
This is when Anthropic and its competitors decide the fate of humanity."
[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]
Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.
2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.
3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.
4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.
5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed): https://t.co/TZOm3LfptY
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
I left Google DeepMind in June. Jacob is right: many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.