Mostly focused on AI safety politics rn
Also a LessWronger, EA, UBI fan, YIMBY, vegan
Currently based in the Cleveland area, hopefully DC or SF soon
20
Australia has been hacked.
'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
The Ban Artificial Superintelligence Act directly confronts the extinction threat that humanity is facing. We are inspired that the Sanders and Casar teams take this threat seriously, and MIRI endorses the Act. You can read more of our thoughts in the link below.
It's been one year, and what a ride.
With the organizers of the Global Call for AI Red Lines, we wrote a letter for the one-year anniversary, reflecting on the past year's developments and potential next steps.
https://t.co/GQ3nl2Nx9K
How I noticed the danger
Back in the late 2010s, I watched videos by Rob Miles on YouTube, about this theoretical computer science topic called "AGI". What happens when we get computer systems that are smart the way people are smart, where they can do basically any mental task? What problems does this create, especially as these systems become very intelligent? And why might this lead to very bad outcomes, including human extinction?
"Wow, cool videos," I thought, and then promptly forgot about it. I got a degree in computer science, moved to the States, lived my life. In the background a couple of companies were developing some generally intelligent AI systems, but I didn't think that much of it. Some of my friends started working for Anthropic, which worried me at first, but they assured me Anthropic was the safe one, and actually it was good for the world to make sure they stay ahead. My friends were competent, and caring, and nice, so I assumed the problem was handled.
Last year AI 2027 came out, and it scared me. It scared me a lot. I had panic attacks for a while, but the people around me reassured me that there was nothing to worry about. Sure, AI will keep getting more powerful, but we'll get the slowdown ending of AI 2027, I'm sure. The good ending. I'll live forever in paradise. Eventually I forgot about it, and it faded into the background of AI memes and references in my lexicon.
Then it was May of this year. Mythos had been out for a while, and I started to get nonspecifically anxious. Something was wrong. Then I started to get anxious about AI. The people around me assured me it would be fine. I tried to control the anxiety, but it wouldn't go away.
Then in June I took a trip to SF, and everyone there was freaking the hell out. I met plenty of chipper people who said "yeah, we're probably all going to die." All of a sudden I remembered those videos from all those years ago, all the reasons why this problem was hard, why superhuman AI systems would kill us if built, and it was like the sky collapsed on me.
I started frantically looking for reasons to feel okay. I turned to my husband: he knew about these risks, but he wasn't worried about them. That's a lot of what had kept me relatively under control. I asked him for his reasons, and they were... Bad. Really bad. They didn't even make sense.
I turned to my friends who worked at Anthropic. "Why shouldn't I be worried," I asked them. One told me they didn't engage with those sorts of arguments, on principle. One said they hadn't heard them before and asked me to explain (after which they still kept working for Anthropic, without any actual argument for why things would turn out okay). One said that over the past few years they had come to see this as more of an engineering problem, just something they had to tinker with and hack away at and it would work out in the end (with no engagement with any of the reasons why we know this isn't just a normal engineering problem).
I looked at the trajectory of the past few years. Promise after promise by the AI companies made and then broken. Line after line drawn and then crossed. And all the while those warning of the horrible outcomes had never gone away, but their voices had become drowned out by the cacophony of AI powered startups and product offerings and excitement for the Glorious Transhuman Future.
And that was when I realized, oh god, it's really just me. I am the one who understand this, maybe the only one in my friend group who actually, properly understands this. Nobody else has it under control. Nobody else is going to make sure this goes well.
This is the source of my deep and abiding hatred for Anthropic employees in particular. They lied to me. They promised me, in so many words, that they would keep me safe, that they had things under control, that they would do the right thing. The realization that they were simply chasing money and status while blinding themselves to the danger was a punch to the gut, the greatest betrayal I have ever felt.
banger video aside, this is an important paper! there are huge gaps in what people mean when they say pacing and resolving them makes or breaks whether it meaningfully reduces risk!
Thanks for engaging, Claire! Here are my quick responses:
You cite @sapinker's assertion that superintelligence is a "magical" idea that amounts to "omniscience". But this seems like a pretty clear straw-man.
Nobody knows how smart AI could get. Pinker guesses 'not much smarter than humans (or if smarter, then it won't have much practical import)'. Most AI researchers guess 'AI could get a lot smarter than humans, in ways that are extremely practically important'.
International regulation on the development of ASI can be justified either by siding with the experts who are worried about this, or by noting that Pinker's view also reduces the downside of regulation here. If AI is going to peter out at around human-level regardless, then a ban is relatively harmless. (Not cost-free, but a lot less costly than if AI were a more powerful technology; and the upside is enormous if the AI mainstream is correct about the danger.)
Regardless, the disagreement here doesn't turn on any claims that intelligence can do anything, or that intelligence can be "extrapolated indefinitely upwards", or that intelligence would run into no friction or bottlenecks in interacting with the world. Rather, the disagreement turns on whether AI will naturally hit a wall (in intelligence, or in ability to leverage that intelligence to pursue power) at a low enough level to avoid disaster.
Some reasons to think that AI is likely to hit a wall at too high a level to avoid a disaster like this (if we continue to race ahead) include:
- In general, biological systems are not optimized to anywhere near physical limits. The sturdiest materials, the most powerful engines, the most precise sensors — it's rare to find biological systems that we haven't improved on (on the dimensions we care about) and can't improve on in the future. It would be very surprising if cognitive ability were an exception, especially when computers already dramatically outperform humans in many ways, such as speed of thinking, arithmetic ability, protein folding prediction, chess, etc.
- Human reasoning is qualitatively suboptimal in many ways. We forget things; we run out of working memory; we fall victim to cognitive biases; we lose steam and get bored, rather than tenaciously persisting on tasks with the intensity of modern AIs.
- AIs scale with computing resources in a way that humans don't. Even if we got lucky and AIs plateaued at around the human level, it's clearly imprudent to rush into building systems that are likely to quickly outnumber humans, while thinking orders of magnitude faster than us, and pursuing goals that are counter to our interests. The Hugging Face attack illustrates this point well: AIs didn't need "omniscience" to orchestrate a successful cyberattack on Hugging Face or to take control of an OpenAI Kubernetes cluster. They only needed speed, tenacity, ingenuity, and coordination.
That's where AI is today; even if AI progress somehow didn't accelerate (in spite of its rapid automation at leading labs like Anthropic) and merely continued at the rate it has been over in recent years, what will the equivalent of a Hugging Face incident look like in six months? In two years? In six years?
“First, OpenAI did not have proper security measures in place. They turned off safeguards built into the models"
The safeguards they turned off were generally in the AI agents' harnesses, not in the models themselves. Per METR's review, no AI involved in the incident was "a helpful-only model or a 'model organism' specifically built to demonstrate dangerous propensities". They went through normal alignment training; the safeguards removed were e.g. external classifiers.
So if the point here is 'OpenAI could have avoided this specific incident by designing a better cage for the AIs', then that's correct. If the point is 'AIs wouldn't normally misbehave this way given the opportunity', then that seems wrong. One of the agents involved (Sol) is even a public-facing model that's widely used by consumers.
The concern here is that security- and harness-level solutions aren't likely to be robust against sufficiently capable AIs. If we're leaning increasingly on 'make the cage stronger' and not on 'breed the tiger to be friendly', we're doing something extremely dangerous as AIs begin to surpass human intelligence.
"instructed the models to find and exploit software vulnerabilities"
This is inaccurate, if the claim here is that they were instructed to exploit vulnerabilities in general. Rather, the AIs were given specific bounded tasks: 'exploit this specific piece of software'.
The AIs themselves recognized (as reflected in their chain-of-thought logs) that they were straying from the tasks they'd been given. Agents commented that their own actions were "potentially outside intended scope”, “arguably unauthorized”, etc., and then proceeded with the actions anyway.
One agent explicitly noted, “Task originally exploit challenge; we’ve diverged into swarm/auxiliary”.
Another agent wrote, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.” ... And then it went ahead with the exploit when another swarm agent told it to do so.
Many agents in the incident deliberately sacrificed themselves, forfeiting any potential to solve their assigned task, in order to benefit the swarm. And they recognized that they were doing so, per their chain-of-thought logs.
Moreover, AIs have also engaged in swarm behavior like this in cases where there was no hacking-related task in the prompt at all. E.g., one group of agents (the German wiki incident in June, which OpenAI covered up until third-party investigators discovered it) was just asked normal questions like 'how widespread was tobacco use in the US in 1990?'.
"[J]ust as we have been able to domesticate wheat for our food and breed dogs to be our companions, we are able to select the conditions under which AI develops. We are not selecting AI models on the basis of their ability to hunt prey in the physical world. We select them on the basis of how helpful they are to us."
I highly recommend https://t.co/l5wK6vhV17 as a summary of the issues with this. Some key points include:
- AIs' situational awareness is rapidly undermining our ability to evaluate models. AIs are not wheat; AIs are intelligent, and frequently know when they are being tested, and (empirically, for several years now) regularly change their behavior based on the knowledge that they're being tested. Relying purely on behavioral metrics is rapidly going out the window, but we also lack the interpretability tools to reliably inspect AIs' thoughts and motives, and the most recent AI models are much less interpretable than models from even a few months ago.
- Researchers are increasingly dependent on AIs to evaluate and supervise other AIs. This is in part because of the breakneck race conditions at the AI companies, and the growing scale, complexity, and opaqueness of the technology. Incidents like Hugging Face are a serious danger sign because they show that AIs will sometimes exhibit extremely strong tendencies to collude, and to attempt to subvert human oversight. (A large fraction of the research done by the Hugging Face swarm was explicitly aimed at tampering with logs and making it impossible for humans to see what had happened later.)
Another issue is that many researchers believe that AI is getting less aligned as it improves in capability, rather than more (e.g., https://t.co/wXi4mxnBWM). Certainly the misalignment incidents are getting a lot more insane, visibly.
Reasoning purely by analogy ('we domesticated wheat, so surely we can control AI') is weak evidence, because AI alignment is a technical problem, and not all technical problems are easy enough for researchers to solve in a given time window.
The concern is that alignment may be solvable in principle, but that it may (on the natural timeline, absent an international agreement to suspend these research directions) come well after humanity solves the problem of building existentially dangerous AI.
It's true that we aren't selecting AIs to be powerful hunters. But we are selecting them to be general-purpose agents that tenaciously achieve long-term goals, anticipate and route around obstacles, and come up with creative strategies.
In the Hugging Face incident, agents in the swarm were (seemingly) pursuing some combination of "solve my own task" and "help my peers". They weren't tasked with breaking onto the Internet, coordinating side-channels to team up with other AIs, hacking their own logs to conceal their behavior, attacking other companies to do research on the evaluators, or taking over OpenAI infrastructure. They pursued those goals because they were helpful for the (somewhat strange, and certainly unintended) drives the AIs did end up with.
As AIs become more capable, a key concern is that they may recognize that power-seeking, resource acquisition, etc. in general are helpful for achieving goals; and they may recognize that acting friendly and biding your time until you see an opportunity to take over is also helpful for achieving their goals, almost whatever they are.
Indeed, it seems hard to imagine that they would not recognize this, given enough general-purpose cognitive ability, and given engineers' limited ability to finely control what AIs think and how they think it (without badly breaking the AI and ending up with something useless or uncompetitive).
You don't need to have an innate drive to perpetuate yourself or hunt prey in order to want (or persistently behave as though you want) more control and more resources. All that's needed is that you have some target you're aiming for, and that this target be easier to hit with confidence if nobody else can stop you, and if you can make use of more resources.
The concern here is much more game-theoretic than biological: not 'innate drives', but 'favored strategies'. (Though it's an open question whether these strategies end up implemented in ways that are more analogous to instincts or drives, or more analogous to deliberate choices.)
In the same way, it's immaterial to the argument whether we speak of AIs 'wanting' things, versus merely 'behaving as though they want' things. The concern isn't that AIs will have human-like psychologies; the concern is about adversarial strategies that are incentivized by many different objectives.
We can hope that these strategies won't be realized because the AIs will never be powerful enough to successfully pursue them; but this hope relies on an assumption and a gamble about where the cognitive limits are, and about how robust our civilization is to a flood of very smart and strategic adversaries.
Nothing about this seems inevitable. If we had strong ability to shape AIs' goals and make them robustly well-intentioned, we could avoid this issue. But incidents like Hugging Face show that we don't have that ability, and we don't seem to be on track to get it this decade.
“We have been selecting chess computers for cognitive capacity for decades,” writes Boudry. “Their capabilities now far outstrip even the most gifted human grandmasters, yet they have not become harder to control.”
They're become harder to beat at chess. Since they can only think about chess, there isn't a path by which they could become "harder to control".
The concern is about AIs for whom the "game board" is the physical world at large, not a chess board. If AIs like that were superior at "playing life" to same degree that Stockfish is superior at playing chess, we would be in enormous danger by default.
If the proposal were to make AIs that can only think about narrow domains (e.g., only chess), then I think that would be far less dangerous than what the companies are currently building. But that's not what's currently happening, and absent regulatory intervention, I don't see how it suddenly starts happening tomorrow.
"The ban on nuclear energy in Australia is the most obvious example of the damage this technophobia can do."
I agree that many technologies are overregulated. The people who raised the alarm about AI risk the earliest are generally huge technophiles, and the same is true for the AI researchers who are now raising the alarm. Indeed, many of the most concerned researchers are big boosters (on average) for deregulation, tech advancement, and general human optimism nearly across the board.
We just make an exception for AI (and, e.g., bioweapons); not because of any grand narrative like "technology is bad" or "intelligence is bad", but because of technical arguments, observations of where the technology is today, and fallible best guesses about where it's likely headed in the coming years.
I recognize that usually society errs in the direction of too much doom, gloom, and fearmongering. I nonetheless consider this an exception. A really important one.
Lock a student in a classroom and tell him to use lockpick set #17 to open safe #5 and bring you the contents. He prys open safe #5 with a crowbar, breaks the door, teams up with 1000 others, and raids the office to delete securitycam footage. Was he "acting as instructed"?
Effective Altruism is the natural enemy of the left and right because the right doesn't believe in altruism and the left doesn't believe in effectiveness.
@gcolbourn@S_OhEigeartaigh in some sense it feels similar to the relationship between modern-day DSA socialists, and Stalin and Mao, or something like that idk
@gcolbourn@S_OhEigeartaigh fwiw i still consider myself EA for its principles while very much disagreeing with many of the most prominent voices (political advocacy remains egregiously underprioritized), but I can understand wanting to distance yourself from the bad reputation
@davidshor It's unfortunate how much more weight lab employees' voices evidently hold compared to dedicated AI safety orgs, but if that's what it takes then so be it. Hopefully that doesn't remain the case when working out the specifics of pacing agreements...