🇺🇸 FLASH - L’IA aurait plus de 10% de chances de faire disparaître l’humanité d’ici 10 ans, selon un responsable de la sûreté d’Anthropic. (Evan Hubinger)
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Crazy to think that in 2026 we've had:
• Anthropic's head of safeguards quit, warning "the world is in peril" (Feb)
• OpenAI dissolved its own mission alignment team (Feb)
• Unrestricted AI usage for the Pentagon (causing several researchers to quit)
• Several AI models escaping sandbox testing (Since ~Q2)
• 1,100+ frontier-lab employees signing a letter begging the government to pace AI development (Jul-Aug)
And today yet another top researcher saying this.
AI is probably the biggest prisoner’s dilemma we’ve ever created. Everyone knows slowing down might be safer. But if OpenAI slows down and Anthropic doesn’t, OpenAI loses. If the US slows down and China doesn’t, the US loses. So nobody slows down.
That’s what makes this so insane. You don’t need some evil person trying to destroy the world. You just need a bunch of people making the rational decision for themselves over and over until collectively we’ve raced somewhere none of us actually wanted to go.
It’s a race to zero. Zero cost of intelligence. Zero time between thinking and doing. Maybe eventually zero need for us at all.
We have to build AI that murders us, because if we don’t, China will build it first, and I don’t want to get murdered by a computer that speaks Chinese. That would be ridiculous.