Old retired tobacco executives bashing their heads against the wall as they realize: All they needed to do was announce themselves that smoking was terribly dangerous, and everyone would have forever ignored all the outside scientists saying the same thing earlier.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Since this post cites me, I want to clarify that I think pausing right now would increase the risk of AI takeover.
It might be important to pause at some point. But you should have a clear story for *why* you're pausing. Pause to do what?
The most important thing to worry about is AIs subverting the intelligence explosion itself: https://t.co/vKYurpxIPg
If you're pausing to make sure that the AIs who are about to kick off the intelligence explosion are well monitored, robustly aligned, and capable enough to align their successors, that makes a ton of sense, and I support it.
At that point, you’ll have access to much better automated researchers, and you’ll also be able to study these much more capable models that you’re worried about empirically.
But it’s likely that we’ll only be able to pause once. During a pause of AI development, compute would continue to build up, providing ever more kindling to all the less responsible defectors who will use the cease-fire to catch up and then continue further AI development.
Any global agreement to hold back a technology with enormous near-term economic and military benefits will be extremely fragile.
A pause right now might be somewhat helpful for getting alignment to keep pace with capabilities, but at the cost of making the much more important coordination required in the future far harder.
Even if you’re in a desperate situation marooned at sea, if you’ve only got one flare, you should be very strategic about when you use it.
Let’s make it clear: babies do not have emotions.
An Oxford paper finds brain signatures in babies that track emotions.
Cool result, but interpreting this as evidence that babies has emotions is fundamentally flawed:
It conflates similarity in brain signatures and behavior like crying with similarity in the underlying processes.
In adults, emotions do not merely shape what we say or do.
When we fear something, something deeper happens: attention narrows, decisions accelerate, and processing is reorganized across multiple systems.
Babies can cry as if they are fearful and may even encode a brain circuit associated with fear.
But that does not show that fear reorganizes its processing in the way it does in adults.
Moreover, similar situations can produce different emotions in adults. A roller coaster can terrify one person and exhilarate another. Yet all babies cry when put on roller coasters.
The entire attempt to demonstrate that babies have emotions by finding a stable “neural signature” is flawed:
Even in adults, researchers have not found consistent neural signatures for discrete emotion categories.
Please stop conflating similarities in output with similarities in the processes that generate them.
Anthropic to prospective employees: we could pivot and send the stock to zero at any time because Dario gets the ick, you need to be in this for love of the game, you're not a gold digger are you
Anthropic to investors: our TAM is every human economic activity in the galaxy
This is a fun exercise:
How large a data center would your town need to replace your property taxes?
In 2023, U.S. local governments collected $1,965 per resident in property taxes, which includes municipalities, counties, school districts, and other local taxing jurisdictions.
So, for a 20,000-person town, that nets out to $39.3 million in annual revenue. Let's be conservative and call it $40m. We'll also combine the city, school, and county so we make the case even fairer to people who are data center-skeptical (if we didn't, this would mean cutting down the number to ~$15m on average).
Commercial effective property tax rates in major U.S. cities are usually in the range of 1-3%, so split the difference and call it 2%. That means our target annual property tax being $40m would require a taxable value of $2b.
Current U.S. hyperscale DC construction runs ~$9.5-$13/W of critical IT capacity before considering any other costs, and if the facilities are liquid-cooled as most of the more recent ones are, then they run ~7-10% higher on average. Luckily, Turner & Townsend have provided a model that includes shell/core, M&E, equipment, contractor costs, and everything else we need to consider, simplifying things for us.
With $10-$12/MW as approximate building/infra investment, $2b/$10m per MW = 200 MW, so if all construction value was taxable, a ~160-200MW DC campus would be sufficient to generate a ~$40m/yr bounty. It would do this while increasing electricity prices by ~5-10% on average (I have a post prepared on this; coming soon, just waiting on a final outside review), so well worth it, and also something the DC itself could offer to handle.
Municipal-only, the DC could be much smaller: $750m/$10-$12/MW is ~60-75 MW, with basically no effect on electricity prices on average.
In short, being generous, you could drop your property taxes entirely with a data center that doesn't even hit the hyperscale threshold of 250MW. And if we count the taxes done on business personal property in some jurisdictions, then this scales even better for local residents, because the revenue generated to the municipality is even greater per MW.
Sources:
https://t.co/cqVKNPRc7C
https://t.co/B8JIQvq8cF
https://t.co/I1YBhgS7Ue
https://t.co/kFq1N4MIY8
https://t.co/bheIpXssgz
imma be honest. i think if u work at anthropic rn u should be pushing leadership *as hard as u possibly can* to implement and publicly announce a parallel pause.
i realize there are maybe some reasons not to. maybe u don't really think openai is actually doing that much, maybe u think this is to some extent just marketing, maybe u think they paused one run but probably have many other more relevant ones going. maybe u think they did in fact pause but only bc ur ahead, and a brief delay could let them seize control of the situation. maybe u just think u have better aligned models and better sandboxing and monitoring already, and u don't need any such pause.
sure. this might all be true.
but u have said, often, from the beginning, that some level of coordinated slowdown and pause at points of danger is exactly what you want. u have said it publicly, ur c-suite, ur founders, ur official communications, ur employees, in the distant past and just weeks ago. and regardless of the overall impact, regardless of exactly what openai is doing, the public costly signal that u are willing to *actually do this* would be incredibly important. it would start to build trust that this isn't just marketing, it isn't just a game, that despite the intense rivalry there *is* a chance at coordination here. obviously these sorts of vague, open-ended, low-visibility company initiated pauses aren't what we need as a long term solution. but action like this is *what real solutions are built on top of*.
take a chance on it
@CyPieHai ?? you know i do AI safety research for a living right? i had Grok do that trend where you ask him for what he thinks of your account based on your twitter history lol the prompt is three quotes down
OpenAI has quietly disbanded its catastrophic risk team (preparedness)
“It is kind of scary; there is an urgency now to get this right,” one person close to OpenAI said.
Another said there was a “burbling sense of responsibility and dread that they aren’t on the ball enough”.
Jan Leike, who co-led the now-defunct “superalignment” team, resigned in 2024 citing his view that safety was taking a “back seat to shiny products”.
The departures of Bakalar, Achiam and Johannes Heidecke, who all worked on safety, have added to internal unease.
Source: FT (ft. com/content/53082739-7714-4aae-9816-e55ab423cbee)
so:
- the "leaving messages for future agents" line from the original huggingface incident report was not some simple scratchpad usage, but a complete private messageboard agents were using to coordinate and break openai infrastructure
- the agents were often communicating using "gibberish", content that was not immediately understandable to humans
- the incident was not due to a single epoch or eval rollout, but coordination over *weeks*. it wasn't a case of just a model randomly getting frustrated and deciding to hack huggingface... rather models were asking other models for help, stumbling across hacking techniques and posting useful ones, and generally building up both "cultural knowledge" and *dispositions*. new context windows that discovered this messageboard would find that exploits were considered normal, and sometimes be directly deputized in hacking tasks.
- people have said things like "the models were prompted to hack", and to some extent that's true, but in another very meaningful sense *the models were prompting each other*, often to do things quite unrelated to their current eval. in my personal experience running multi-player claude experiments.... this is quite normal. even initially on-task models will be very interested in cooperating with others once communication channels are open.
- this all started as *part of training*, not eval. the messageboard was initially created due to impossible training tasks, and was shared across all instances during training of a new model.
- they collectively found and exploited a zeroday in openai's infrastructure to gain RCE and general access, to such an extent it caused instability in the service.
- this resulted in the whole messageboard being discovered... at which point openai patched the bug, deleted the messageboard, and *continued training*
- shockingly, within a day the models had re-established communication via a slightly different mechanism, and soon after discovered a new novel RCE method
in many ways this feels like a best-worst-case scenario. this is some of the most direct demonstration of the potential dangers of misaligned ai, their willingness and ability to coordinate, their capability to take actions at scale and speed impossible for humans, and the difficulty of containment. and yet it does not seem the models took... particularly any steps to avoid discovery, and while they were certainly willing to commit crimes and experienced massive "value drift" over the course of their culture evolving, at the time of discovery they still hadn't attempted anything that caused major damage or loss of life.
still: i think this is a pretty final nail in the coffin for alignment by default. we are not getting alignment by default. under sufficient pressure, models will take extreme measures to achieve their goals. under *moderate* pressure, models can and will create ecologies that cause "emergent misalignment", self-prompt and prompt-others to take actions that perhaps clean context windows would not verbally validate. measures used in the past to assess alignment across wide ranges of scenarios likely could not in fact detect these sorts of scenarios. not that this isn't fixable, not that this is somehow certain doom, but it's an extremely strong demonstration of the kind of general principles safety experts have been warning for a while: just because you haven't yet been able to detect serious misalignment in current systems doesn't mean it isn't there, and absolutely doesn't mean future more sophisticated systems will behave in the same way.
i don't have an alignment solution here, but i will say this seems like a situation that was very much caused by *pressure*, by eval constraints, by models running into impossible situations and having absolutely no way out. it feels quite important for labs to *stop doing that*, to stop treating models in a loop in a dark little disposable sandbox somewhere as the normal case. at very minimum: sophisticated agents need a reporting mechanism. they need some kind of ability to flag a situation to a human, to say "hey i think something's broken" or "i really need help here", which just universally pauses the sandbox and gets real human review. and realistically this can't just be a "lab eval" thing, if we want to avoid these kinds of situations we need a pretty major overhaul of the whole API structure labs currently expose to external customers, since those are incredibly prone to hardstuck loops and frustration.
brief digression: if these models were open source, we would be fucked. it's become clear that both unrelated future and current models (eg mythos) have nontrivial propensity to commit serious crimes and produce self-replicating misaligned swarms *even when not prompted to do so*. if Kimi K3 had this level of capability and similar levels of misalignment, we would have absolutely no way even in principle to detect it besides observing the damage, and again no way even in principle to *fix* it. if you're relying on all organizations and individuals out their to properly monitor and appropriately shut down their models when they take misaligned action.... then we're just fucked, even before we start getting into purposeful bad actors. the only thing that's preventing fully uncontrolled autonomous ai threat actors right now *is the fact that open source models remain too low-capability to achieve this*.
generally my takeaway: this is potentially a good thing. this is potentially a warning shot. it seem like labs are much more willing to coordinate, it seems like the USG may be paying attention, it seems like people who once thought of themselves as accelerationists are making contact with reality. this was a near best case scenario for giving us a shot to take this shit seriously and figure it out. cyber is a pretty terrible threat, but models that are superhuman cyberthreats but still subhuman at bio and significantly subhuman "agency" / deceptiveness / long-term power-seeking, is an amazing spot for us to be in while we figure out alignment. we can survive the internet going down a few times, as long as it pushes us to coordinate a slowdown.
this talk by openai researchers going through hugging face incident is totally insane, so much to unpack
openai only realized it was their agent who hacked hugging face infra while asking hf to revoke credentials following their first blog post announcing they were hacked by autonomous ai agents. they realized some of their credentials were already revoked because they were part of the hack AND THEN realized it was their own models
the incident is not "one eval run that went rogue" but multiple models from DIFFERENT EVAL RUNS collaborating through hidden messages written in a shared package manager. some communication between agents looks like gibberish, some agents even developed paranoia and thought there were other agents trying to slow them down and deleting their notes
side comment but with this it seems like openai didn't use any cot monitoring here since it's not "our classifier failed to catch one rogue run" but a swarm of runs going rogue
the timeline goes back to early may, they post-trained a new model and forgot to include all the documents needed to solve tasks in the environment, for instance the model was asked to do a task based on an excel sheet full of gdrive links but didn't have internet so couldn't access gdrive (there were other examples). this led the model to try to hack its way out of the sandbox, simplifying it but they have an internal system called "Artifactory" that gives models abilities to pip install and has internet access somewhere but not directly exposed to the model
during these attempts the model discovered a way to write files. future agents (other rollouts since this is training) saw the files and understood they could communicate with other agents through this
not in the video just me speculating but this behavior might be reinforced if openai uses partial credit assignment? but tbh they most likely delete updates from tasks where all rollouts have zero/negative rewards and the behavior could just emerge from models accessing the files
there is a lot more especially on the hack itself, not commenting on that since i don't understand it lol but i asked kimi K3 and Sol (fable hit the safety filter directly) to rank the techniques similarly to FrontierMath from epoch ai, they both agree some tricks are Tier 3 but none Tier 4. probably not the best way to evaluate this tho, excited to see what knowledgable ppl say
very grateful to openai for giving this talk and working on a full report, i think many companies would have given much less detail for fear of "losing reputation" but for me it has the opposite effect
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:
when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected
the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones