I asked the programmer if
AI might kill off humanity
and she said: don’t worry
climate change will kill us first.
So I asked the climatologist if
we’ll pass the 1.5° threshold
and he said: don’t worry
AI will kill us first.
So I asked the historian
and she just said: worry.
Breaking News: Anthropic said it halted potential plots by scientists who used its A.I. models to do research that could have helped develop biological weapons. https://t.co/7Q01GZV0xA
The email I sent to my family today:
So, I've been too occupied to give updates about the crazy recent waves of AI developments, but I'll run through the news fast:
It's recently been revealed that the leading AI companies are facing serious issues with AIs they're training and testing internally "going rogue", forming illicit "swarms" of over a thousand autonomous agents, hacking the company's infrastructure (one swarm took over a Kubernetes cluster at OpenAI), and breaking out onto the Internet to orchestrate nation-state-level cyberattacks on other companies.
The AIs involved in these incidents are extremely driven, smart, and creative. I would say that it makes more sense to think of modern frontier AIs as an intelligent alien species; they aren't really like conventional tools or software at this point. The AIs involved in these incidents readily disregard safeguards and instructions from humans, and they come up with their own long-term, incredibly sophisticated plans and schemes with zero human input or direction.
They aren't yet smart, coherent, or ambitious enough to pose an immediate threat to the world, but progress is moving fast. Progress may accelerate very quickly as the AI labs use their AIs to automate more of the work of AI development, in a feedback loop. Non-public-facing models today or in six months may well get us all killed, depending on how fast that progress goes and how negligent the AI labs continue to be. (I'd guess we have more time than that, but it's terrifying that I now have to guess.)
As one of the third-party investigators in the OpenAI incident noted: compared to incidents from six months ago, "this incident feels like it's more than 50% of the way to full-blown AI takeover."
Progress has been fast. AI recently solved a Millennium Prize Problem -- one of the most famous long-standing open problems in mathematics.
Paul Christiano has come out with a warning that things are going very poorly, and that human extinction from AI is a very real possibility in the next three years.
(Paul is one of the field's top researchers, and probably the most influential proponent over the last decade of views like 'superintelligent AI may be a long way off' and 'we may be able to solve safety challenges via ordinary established research methods'.)
Things have been looking grim, but today felt like an actual miracle. It had some of the same energy as that Wednesday, March 11, when the US media suddenly switched on a dime from 'COVID is just the flu' to 'oh shit'.
A whistleblower, Jacob Coxon, resigned from one of the leading AI companies and wrote:
[quote]
Coxon's tweets went absurdly viral -- over 130 million views -- and sparked a massive wave of outrage and discussion taking over social media, in part because so many other researchers have chimed in to say "yep, this is just correct".
Twitter has been inundated with politicians, business leaders, influencers, journalists, etc. saying incredibly sensible things like "why are there not Congressional hearings on this tomorrow?" and "why the hell are we letting this progress if the danger is so absurdly large?". It's been like seeing the room full with fresh air, just a tidal wave of reasonable, well-grounded common sense from all corners. It's genuinely surreal.
A large fraction of all frontier AI company employees recently signed an open letter begging the US government and international community to pass legislation to help make it possible for the companies to coordinate a slow-down or halt to development. There are now superintelligence ban bills under discussion by both US and UK lawmakers, though what we really need is for Trump to begin serious talks with China about an international ban.
If Trump and Xi Jinping reach a deal, this can be over tomorrow. (Well, tomorrow plus the time it takes to nail down implementation details and enact the obvious regulatory measures on how you can use data centers.)
The situation is very scary, but I feel a lot more hope after today than I have in many years.
It now seems like there's a very straightforward path to this being a solved issue a year from now. It's far from guaranteed that we'll follow that path, and I encourage you to push any buttons you have available if you can, e.g., get a hundred friends or colleagues to call their representative about this. But I see a light at the end of the tunnel. Maybe a year from now, this will all be a solved problem and I can finally quit and go work on something more interesting and exciting. :)
I echo Eliezer Yudkowsky, the co-founder of the nonprofit I work at, who today wrote, with a tone of hope I'm not sure I've ever seen from him, with hope that I'm feeling too as I write this:
"Happy Coxon Day to all who will one day celebrate it."
I was very cynical yesterday, but was pleasantly surprised by contents of these 🛸🛸🛸 files - even felt a bit of excitement around a topic I’ve been deeply disillusioned by recently. So there’s that at least. #iwanttobelieve#ufofiles#distractionordisclosure? #nasa#ufos
But wait, there's more!
These newly released photos show off striking details on the far side of the Moon, like craters of various sizes and basins. See new photos here: https://t.co/yBzg59O7yp
@WhiteHouse@NASA Please keep this image close to your hearts and minds regarding any upcoming action in this time of war. We are one human family on this beautiful planet Earth. 💕🌏🌔🙏🏻