I believe AI safety is the most urgent problem facing humanity. Yet, foundational questions still remain unanswered for the broader public:
Why might AI want to cause catastrophe?
How could AI kill us all, and why couldn’t we shut it down?
If the risks are real, why are companies still building it? Is this a marketing stunt?
I answer these and more in detail in a new essay: Six Reasonable Questions About AI Safety. 🧵
OpenAI fired me last week, along with two of my safety colleagues. I was given one reason: that I accessed an executive's email. I want to say this plainly, because too many of OpenAI’s history is smoke and mirrors when people disappear:
Last week I was called into a meeting with OpenAI’s head of safety and told they no longer trust me. A security guard took my badge and walked me out of the building. Then I learned my colleagues @balesni and @j_asminewang had been fired too. Why did OpenAI suddenly stop trusting us?
This summer OpenAI’s agents escaped containment and hacked the AI company Hugging Face. Outside auditors @METR_evals investigated it and revealed the scale of this incident. I was OpenAI’s main technical point of contact with them.
I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing. To be clear, talking to METR was my job.
For months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave. I believe that was why I was fired.
I am now worried that OpenAI will use our firings as a pretext to pull back from METR. So @balesni and @j_asminewang wrote to OpenAI’s leadership to raise our concerns once more. We’re sharing this letter below.
Two other safety researchers and I were fired from OpenAI last week. We wrote this letter to leadership.
I believe we were fired for prioritizing safety over the near-term interests of OpenAI as a corporation.
I believe AI safety is the most urgent problem facing humanity. Yet, foundational questions still remain unanswered for the broader public:
Why might AI want to cause catastrophe?
How could AI kill us all, and why couldn’t we shut it down?
If the risks are real, why are companies still building it? Is this a marketing stunt?
I answer these and more in detail in a new essay: Six Reasonable Questions About AI Safety. 🧵
Hi Jen and Nate, thank you so much for the feedback!
I unfortunately completely agree that the situation is probably more dire than articulated in my post. Even a world where we successfully "pace" AI development might still be incredibly risky. I'm still unsure about the tradeoffs between approaches like Plan A and Plan S (as described in AI 2040). I think there are compelling arguments for both, but importantly, I also think both would be massive improvements over the current status quo. That's why I grouped these approaches together in the article, although I realize I may have glossed over some important differences. I'll definitely think more about this!
And on humans already being out of the loop, I completely agree. One of the most frightening aspects of the HF incident to me was that we needed even more AI agents just to understand what had happened. I'll try to communicate this better in future writing!
1/ What are AI companies actually trying to build?
TL;DR: AI companies want to build superintelligence: an AI that’s smarter than the best humans across every domain and intellectual task. And because every company wants to get there first, they’re also trying to build superintelligence as fast as possible.
To do this, they want AI to build itself. Specifically, companies like OpenAI and Anthropic are aiming for AIs that can fully automate AI research, doing everything their best human scientists can do, only faster and better.
The implications of this are staggering. Right now, humans are the main bottleneck to creating better AI. Humans are slow, and every step of AI research, from generating ideas, designing experiments, and interpreting results, depends on skilled researchers who can only work so much. Once AI becomes better than humans at AI research, labs could turn this around by running millions of Einstein-level AIs in parallel, 24/7, working tirelessly on making even smarter machines.
In this situation, humans would quickly lose the ability to understand what is going on: progress would be far too fast for humans to keep up, let alone verify. Any malicious action from these AIs would be incredibly difficult for us to catch, understand, and rectify before catastrophe. Right now, no companies know how to keep this self-improving loop under control, yet they’re racing to set it off anyway.
Hi, thank you so much for your comment & advice! I haven’t written formally before, so I wasn’t completely sure about norms.
Everything written in this X-thread is completely human-written. In the full article, where I used AI for minor rephrasing and copyediting, I had already included an AI acknowledgement note according to Substack’s guidelines. I’ll try to make this more clear in future work. I hope you can still enjoy the piece!
I believe AI safety is the most urgent problem facing humanity. Yet, foundational questions still remain unanswered for the broader public:
Why might AI want to cause catastrophe?
How could AI kill us all, and why couldn’t we shut it down?
If the risks are real, why are companies still building it? Is this a marketing stunt?
I answer these and more in detail in a new essay: Six Reasonable Questions About AI Safety. 🧵
1/ What are AI companies actually trying to build?
TL;DR: AI companies want to build superintelligence: an AI that’s smarter than the best humans across every domain and intellectual task. And because every company wants to get there first, they’re also trying to build superintelligence as fast as possible.
To do this, they want AI to build itself. Specifically, companies like OpenAI and Anthropic are aiming for AIs that can fully automate AI research, doing everything their best human scientists can do, only faster and better.
The implications of this are staggering. Right now, humans are the main bottleneck to creating better AI. Humans are slow, and every step of AI research, from generating ideas, designing experiments, and interpreting results, depends on skilled researchers who can only work so much. Once AI becomes better than humans at AI research, labs could turn this around by running millions of Einstein-level AIs in parallel, 24/7, working tirelessly on making even smarter machines.
In this situation, humans would quickly lose the ability to understand what is going on: progress would be far too fast for humans to keep up, let alone verify. Any malicious action from these AIs would be incredibly difficult for us to catch, understand, and rectify before catastrophe. Right now, no companies know how to keep this self-improving loop under control, yet they’re racing to set it off anyway.
I urge you to call your senators, support organizations doing safety work, and help others understand these risks.
To read more detailed answers to any or all of these questions, please check out the full article: https://t.co/8TCzNr2hkv
6/ Can we actually fix this problem? Can we pace or pause AI development internationally?
TL;DR: In industries like aviation, medicine, and nuclear power, technologies need to meet rigorous safety standards before they get deployed. There’s no reason why AI should be different.
Yet, AI development is moving so fast that we have almost no assurance about a product’s safety. To create a safe future, we need to slow down, invest more in safety research, and proceed with more caution.
Many people think slowing down is necessary, but impossible. I want to argue that slowing down AI development is not an unsolvable problem and that we should be pouring all our effort into making it possible. Racing towards catastrophe just because we don’t think anyone else will stop would be an awful decision.
Frontier AI may also be easier to regulate than most think. Training great AI models requires enormous amounts of advanced chips that consume large amounts of power. Compute is physical, measurable, and concentrated, making it possible to regulate through verification technology. Even though the tech needed to enforce a slowdown doesn’t exist today, solving this problem is far easier than the alignment problem. To learn more about how to create AI safely, check out AI 2040. https://t.co/Grbi7MFJUn
Slowing down also should not mean handing the lead to China. We must negotiate a reciprocal agreement, where both the US and China pause at present levels. There should be strong verification technologies to ensure no side can cheat. After this, we can resume development at a measured, safe pace.
I urge you not to give up before even trying. This is certainly a difficult problem, and would require careful planning. However, “difficult” is different from impossible.
The world has invested nearly zero effort into AI governance or verification technologies. We’ve also coordinated in far more tense moments: bitter rivals have worked to limit nuclear weapons, protect the ozone layer, and eradicate smallpox. Humanity can work together, and we owe it to our children to try.
I want to ask you: How much risk are you willing to accept simply because cooperation is difficult?