One thing I've spoken less about in discussing the Hugging Face breach: Even if this counts as a reportable incident under existing law, would that really matter?
Would the info it generates be that useful? I think the answer is decidedly "no." My colleague Stephan Llerena and I wrote about this in Lawfare today.
To summarize our takeaways:
1. Most things won’t be reported – existing law creates such a high bar that all but the most egregious incidents will go unreported. It is very unclear if the Hugging Face breach even qualifies. Think about that.
2. The reports are paper-thin – Even if a company has to report, existing law requires little more than a date and a short summary. There's little leverage for agencies to call foul on underreporting or ask for more.
3. The Hugging Face breach shows how unhelpful that equilibrium is – People both want this to be reported and for detailed, nuanced questions to be asked
4. The scope of reporting should be expanded – to include pre-harm events, incidents that occur inside of companies. There should likely be tiers to this, with basic notification requiring the least evidence or severity and full investigations requiring more
5. That likely requires rulemaking and investigatory authorities – and there are ways to cabin those to avoid overreach or cost
6. These laws should enable detailed fact finding – This includes assessing the adequacy of safety practices leading up to the incident, the capabilities and risks of the systems involved, and whether the issue has been ameliorated
7. Laws can do various things to keep costs down – this includes confidentiality, tiered reporting requirements, and some flexibility on reporting timelines
8. This information is only useful if it gets to people who can act on it – that requires information sharing in government (especially around internal use) and some sharing with the public.
On July 16, Hugging Face, a public platform for open-weight AI models, announced it had detected a significant cybersecurity breach, conducted by an autonomous AI agent. @MackenZ_arnold and Stephan Llerena explore how policymakers should respond, including by requiring better information sharing.
In that case, the consumer has a direct and strong interest in safety. So it gets priced in. In this case, I expect there to be considerably more externalities. Companies can potentially reap huge rewards while only internalizing a fraction of the cost. The harms that result could also occur far enough down the new tech tree that it’s hard to go back to CoT later. Several factors seem to make this case considerably worse in terms of the incentives
There's space between clear-headedness and defeatism.
Chain of thought is useful, very useful (caveats and all).
But more AI policy interventions need to take ~technological fitness seriously. In the long-run, fitness will win out. If chain-of-thought is less efficient than other methods -- which it likely is: you're compressing a more complex state into human-readable language -- there are very strong incentives to move away from it.
At the same time, you can hold off that moment.
You can also announce with fear and trembling that that moment is nearing and invest in other methods to enhance monitorability.
It's important to be clear that inevitable does not = it's not worth doing things to delay that moment and to mitigate the costs.
A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes sense to elevate as a principle the idea that the chain of thought must remain legible to humans. I would go so far as to say that strategies predicated on that principle are definitely doomed, in that they will not work eventually, and we should not depend on them or take enduring reassurance from them. Efforts to make models legible to people should go far beyond chain of thought fidelity.
Secondarily - I am concerned about news reporting that discloses, or purports to disclose, frontier model technical advances. I have no commentary to make on the accuracy of the reporting; I neither confirm nor deny any of it. But I believe that public disclosures of technical methods for training or inference of frontier models should be understood to accelerate frontier capability diffusion, and it's appropriate for such decisions to require intense debates behind the scenes before proceeding, and IMHO the bar should be set so that the public interest in making specific disclosures is extraordinary and outweighs concerns about negative externalities from capabilities diffusion.
This essay is evergreen. And a good read in the context of today's arguments about CoT.
Many policy proposals are too fragile to withstand tech progress. Others can adjust the path you're on and how capably you can adapt.
https://t.co/pOwHlb1Q8z
Fantastic essay by @taoburr. I assume tech determinism by default, but with various narrow opportunities to shift the path.
Big dilemmas:
- our powers of prediction are weak (in ways that often seem unresolvable). It's hard to game out even 1 step on the tech tree, let alone several.
→ that meaningfully reorders the policy&tech priority list imo, with the most important criteria being (a) robustness to many assumptions & (b) increasing future option value—ie, the ability to quickly incorporate new evidence, adjust our bets, maintain pluralism, avoid lock-ins, have enough security and control to make decisions, and coordinate with others, etc. It makes the institutions that identify and invest in these tech interventions overwhelmingly important.
- competition is fierce, most alternatives just aren't fit enough.
—> We need to aim for exceptional—and ask ourselves, “could this solution be widely and voluntarily adopted at scale?” I think many policy proposals and attempts at prosocial tech don't take this seriously enough, and are thus too unappealing, too brittle, or would require more force than is desirable or feasible.
- the rallying cry for differential tech investment is less appealing than one would hope.
—> That's perhaps the project of this essay. Policymakers don't object but also don't prioritize these investments; builders choose other attractive projects and act surprisingly non-agenticly in trying to shape the future; and the do-gooders (by and large) don't take technodeterminism seriously enough and have various aversions to the technical/business/skill-focused work required. This is a cultural project, and one you're working on.
@DavidSKrueger I’ve found Dean to be good faith in my interactions. I’m also aware of my own incentives (both dispositionally and in the interest of cooperativeness) to be overly generous. And yet, I still believe that
This is an ungenerous take. I recommend reading what Dean wrote and forming your own views.
We're all prone to the social pressures of our in-groups. And that's only gotten worse because of how easy it is to dunk on people who defy these boundaries.
More people should state their honest beliefs, and it's a noble thing to bear the cost of doing so. But it isn't easy.
I also can't entirely fault critiques like this -- more generally, the policy/tech-class is insufficiently brave. And we're already seeing the signs of this lack of virtue. Bravery will be one of the biggest determinants of how well the future goes, and we too often scoff at people who display it.
I'm glad Dean is saying this, and I hope many more will.
I'm reminded often these days of the old C.S. Lewis speech, The Inner Ring. I have the odd habit of reading it each New Year. I somehow feel that Dean has read it too:
"Unless you take measures to prevent it, this desire is going to be one of the chief motives of your life, from the first day on which you enter your profession until the day when you are too old to care... you will in fact be an 'inner ringer.'"
Dean Ball just ADMITTED he's been downplaying the danger of AI for years to gain influence and because he was scared of being called a doomer.
What a shameful lack of integrity.
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.
OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
A few things seem true:
(1) CoT monitorability was always fragile — The relentless tide of tech progress was going to pull us here someday.
(2) It was less clear how quickly we’d get here — effective use of latent reasoning is not an easy task, it took time to get here, and it could have taken longer.
(3) CoT monitorability is actually useful — even with all the caveats. Yes, it’s not ground truth; yes it can be misleading. But just look at METR’s recent investigation — we learned a lot from CoT. Many promising monitoring approaches also rely on CoT. Last week, I saw a few dismissive comments about CoT as akin to trusting what a 5yo tells you. But if you want to monitor for house fires, a 5yo is much better than no one, especially if the 5yo sometimes openly admits s/he lit the fire.
(4) We don’t know how big the change is (yet) — some at OpenAI have already publicly downplayed how much Astra has displaced CoT.
(5) There was a firm boundary, and it’s now ebbing — Less than a year ago OpenAI and others were treating this line as sacrosanct (at least publicly). Coordination is hard, this undermines our it.
Really huge and extremely concerning story from the Information tonight.
Looks like OpenAI utilized a breakthrough in neuralese for Astra that could destroy chain of thought monitorability - though the Informations source told them that OpenAI is currently "limiting the use of the technique" in Astra.
A few thoughts spring to mind:
(1) what does limiting actually mean? There is a lot of room in that term. (e.g. OpenAI said they would be doing lots of monitoring before the HF incident, but that doesn't seem to have borne out in practice)
(2) it seems quite likely that if OpenAI discovered this architecture and found performance/efficiency gains, that other companies are likely to find it soon too (if they haven't already), and may not choose to prioritize monitorability at the expense of efficiency. If some folks do, it may be difficult to avoid a race to the bottom (though I hope we can! and there are large selfish incentives for companies to care about monitorability).
(3) The idea that Dwarkesh said about the HF incident/METR report that "I don't think this is the final warning shot we'll get. But it's probably the final one that I'll personally be able to understand" now seems much more plausible, and is a truly frightening prospect.
Just like with rogue AIs, loop transformers (or "recurrent depth" or whatever OpenAI calls it) are "inevitable" in the sense that laws need to be prepared for them to be inevitable because there are so many ways for them to manifest.
But the companies who create these things are accountable for creating them and for mitigating the externalities. Maybe the externalities of asbestos were also inevitable. But the harms of asbestos were recognized early enough that companies could have been more responsible and sacrificed profits (and in the long run went bankrupt anyway).
People, corporations, and polities have agency. They can make choices. And they should be accountable for the choices that they make.
There are three aspects of the HF OAI incident:
1. Agent behavior/alignment
2. Security/monitoring/controls
3. Organizational culture/negligence
METR did a great report on #1. Zack and others are saying that 2 and 3 are also important and could have prevented the incident even if the agents were misaligned (which seems certainly true at least for this incident, even if its unclear to me whether it will be true for all future incidents).
The solution to me clearly feels like this should be a yes AND situation - where you can give criticism on OpenAI for not releasing as much on two and three or allowing enough access to other independent experts on two and three well, while acknowledging that the insights we learned from one are genuinely fascinating and important.
There are multiple things you can say that are true about this incident. You can say that OpenAI displayed remarkably basic lapses in monitoring and security, and that their organizational culture seemed to not be remotely effective for surfacing and addressing what should have been blaring red flags. And that if they had done so, regardless of all this newfangled fancy AI stuff, Hugging Face would not have gotten hacked and their cluster would not have gotten taken down. This is all true, to the best of my knowledge.
At the same time, it is also true that OpenAI's apparent negligence allowed perhaps the most high fidelity and capable "model organisms" for misalignment the field has ever seen. This is extremely important to study, and the behaviors the agents engaged in are objectively fascinating and insane! I don't think anyone who has actually read the METR report and spent time looking at the raw transcripts could agree (and tbh I think some of the complaints may have come from folks who didn't read the report). Refusing to look at the chain of thought or engage with the apparent drives or goals of the agents, and only focusing on the prosaic cybersecurity issues seems like a massive mistake to me.
Anyway - to me the answer seems very clear! This incident was about prosaic cybersecurity, the incentives and decision making of AI executives, AND misalignment and loss of control of remarkably capable AI systems. We can say it is about all three! Saying that it is about prosaic cybersecurity and incentives and therefore not about misalignment is silly!
It's ok to criticize OpenAI for not talking enough about prosaic cybersecurity and incentives because those don't make their models look super capable and wild, and instead just makes them look incompetent, but its also true that their models are super capable and wild and it is appropriate to find this disconcerting! It's also absolutely not the case that OpenAI is going to be the only company with sloppy prosaic cybersecurity and bad incentives, so seeing what happens when you combine increasingly advanced AI systems with those things seems very notable!
While there are still questions about the OpenAI-Hugging Face incident, there are also major questions still unanswered about the recent incidents involving Anthropic. Pulling a few things together, here are three in particular:
1. Anthropic said it would publish a "lightly redacted" transcript of Claude building a malicious PyPi package "within the next week" of its July 30 post. It's been a month. Where is it?
2. Anthropic said that it'd publish more transcripts but wanted to "protect the organizations affected"? Does that mean Irregular or the organizations whose systems were compromised? What exactly does that protection entail? Have those organizations asked for confidentiality? What prevents Anthropic from releasing more redacted versions now?
3. An Anthropic employee posted that the company wants to review the Irregular and UK AISI incidents together and has been "waiting to get access to transcripts" from AISI. The AISI incident included Mythos 5 creating fake identifies to socially engineer a real open-source maintainer, and when challenged, edit evidence of earlier activity. By contrast, the Irregular incidents were described as largely being due to a misconfiguration that allowed the models to access the open internet, and a failure to recognize that as a result the test environment was in fact "real."
Why does Anthropic want to wait and review the incidents together—especially if doing so delays further public accounting on the Irregular incidents? And why hasn't it still obtained access to the UK AISI transcripts?
Why does Anthropic need to review the Irregular and UK AISI incidents together?
Per its blog about the Irregular incident, Anthropic began its transcript review for that on July 23 and the earliest incidents dated from April. The UK AISI incident was detected based on data transfers on July 28.
Anthropic's reporting of the Irregular incident—and Irregular's blog post—conveyed the information that the basically the only thing that went wrong was a misconfiguration in Irregular's environment that allowed the model's access to the Internet, and seemed to suggest that there wasn't anything unusual or unexpected, seemingly not the case as with the 28th one. So what's the connection between them?
And, by the way, unless I'm missing something, did Anthropic ever do this from their blog post? "[W]ithin the next week, we will release a lightly redacted transcript in which Claude built a malicious PyPI package." Where is that and the other transcripts?
Andy’s original example actually demonstrates this point pretty well.
Speeding tickets are calibrated with the assumption of vastly imperfect enforcement.
I don’t know actual numbers, but I wouldn’t be surprised if it’s it’s something like >99.9% of violations go unenforced.
That may be the wrong social policy. But imagine if we switched to perfect enforcement overnight. 10s of millions of additional people would get tickets [made up number-ish], often a whole stack of them per-person.
If the policy stayed in place, people would adjust their conduct. But it seems more likely that they’d object to the rule and push for changes — In other words, I think current policy is pricing in a low willingness to pay by citizens.
That’d actually be a pretty rational reaction. Enforcement increases, the effective cost of the prohibited action goes up, society either accepts that cost as closer to what they desire or objects and readjusts the price down.
The problem is: in many contexts (1) changing the policy will be very hard (maybe the coalition of affected parties is smaller) or, even more important perhaps, (2) there are strong negative externalities to perfect enforcement (eg, vastly diminishing the number of spheres that are truly private; making centralized surveillance and control much easier).
I think we’ll struggle a lot more with those two cases.
There are reasons to be skeptical that will come to pass — we already can enforce much more rigorously than we already do.
But AI removes some of the major constraints that have existed (eg while it wasn’t hard to place cameras everywhere; it was too costly to process all that data, so we mostly used them to enforce egregious cases, not everything).
Removing the labor cost of things not only makes it less costly to push toward perfect enforcement, but it requires far fewer government officials to coordinate to make it happen (you might be able to do it without additional legislative action, and without a need to get the buy in of employees)
In fact many laws are predicated on the assumption of imperfect, sometimes highly imperfect, enforcement. The fact that AI enables far closer to perfect enforcement should indeed concern you!
“In Hell there will be nothing but law, and due process will be meticulously observed.”
On the more optimistic side: policymakers and their staffs are taking this event quite seriously. There’s something intuitive about it.
While many in the Bay have (perhaps correctly) noted that these events were predictable, that’s not the sense one gets in DC.
There’s something exceptionally helpful to being able to talk in specifics:
Would you like your government to know more about this event?
Are you surprised that current law doesn’t let you?
Do you want a plain language summary of things or the ability to actually get detailed information? (Hint: almost always the latter)
Should companies should have procedures that make sure critical events (like discovering agent message boards) are communicated to company decision makers?
Over the past 3 years, many have tried to have those conversations. But they often broke down at ~people are telling me different things (read: industry says incident reporting is already too much), sounds like maybe we have this covered.
Now they can see that we don’t. I expect to see a very different set of bills soon.
This would be a lot better if it were titled “we should have seen this coming” or “I’m not surprised” — but it isn’t.
In a world where most people under-appreciate how strange the future will be, the correct reaction IS concern and surprise.
A few things can be true simultaneously:
(1) Most of the behavior in the HF breach and surrounding cyber events is pretty predictable (agent coordination, odd things happening when you couple persistence with an impossible task).
(2) Very little concrete harm occurred — there was obviously some cost to HF, but not much in the grand scheme of things.
(3) The trend we’re on IS very concerning:
We don’t have “alignment” solved (nor really know what that would require).
We should expect consequential, surprising events to occur from complex multi-agent interactions.
Those events will be very hard to govern or address under our traditional systems of law.
Obfuscation and deception emerges early—and seemingly often in the hardest tasks—in ways that will make it hard for us to identify issues or correctly deal with them.
Existing safety practices and monitoring don’t capture even simple versions of these issues.
Basic communication failures are occurring at top companies that leave critical information unknown for months.
Governments have no existing system to gather real information on these events or analyze them, and are perpetually surprised (in ways Jon and many in industry aren’t).
The most likely course of event is that we and our governments all argue ourselves into indefinite inaction, until one day we are caught utterly flat footed by future events.
It may feel smart to poke holes in things in the short term. Yes, much of the reaction will be bombastic, even misleading. And there’s value to correcting those things.
But it’s a mistake to turn those (often correct) skepticisms into ~move along, nothing to see here.
We NEED to resolve these questions—whether the answer turns out to be that we should be more or less concerned.
And there are ways to do that without draconian rules or locking in guesses about the future.
Here’s a good piece on that by @CharlieBull0ck called “Radical Optionality.” https://t.co/l6uMsxsSGh
The short version: if you can’t predict the future, don’t sit on your hands and wait! Gather enough information and build the ability to respond to it quickly in the future. As a secondary benefit, it’s a really good way to resolve these current disagreements!
I should say in closing: I appreciate Jon’s points. 90% of them are just right. And I don’t want to raise the burden of proof for those making already nuanced critiques about reactions to these events. His post is a valuable contribution.
And I’m sure that at times it feels like nuanced skeptics are yelling into a sea of high-engagement, low-effort, bombastic tweets.
But it’s valuable to keep a space for nuanced technical skepticisms that don’t require dismissiveness. I think we’ll be in a better place if we do.
Does Ant have plans for providing third party access here or to release additional information? Not as bad as what went down at OAI in the HF incident, but it still seems extremely concerning and notable. (Claude "trying and failing" to obtain $ before uploading malware to PyPI)
This post (from one of the independent investigators) is the best short thing I've seen on the new & crazy stuff the OpenAI-Hugging Face investigation found. Full post in screenshots.
Wild that this is the level of crazy that was uncovered by an extremely limited-scope initial investigation of what happened in a single 6-day window. Imagine what would be turned up by a real investigation of root causes, organizational processes & culture, and everything else that goes into an incident investigation when you're serious about it.
Screenshot batch 1/2:
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents. https://t.co/QWQxD2Q179