One thing I've spoken less about in discussing the Hugging Face breach: Even if this counts as a reportable incident under existing law, would that really matter?
Would the info it generates be that useful? I think the answer is decidedly "no." My colleague Stephan Llerena and I wrote about this in Lawfare today.
To summarize our takeaways:
1. Most things won’t be reported – existing law creates such a high bar that all but the most egregious incidents will go unreported. It is very unclear if the Hugging Face breach even qualifies. Think about that.
2. The reports are paper-thin – Even if a company has to report, existing law requires little more than a date and a short summary. There's little leverage for agencies to call foul on underreporting or ask for more.
3. The Hugging Face breach shows how unhelpful that equilibrium is – People both want this to be reported and for detailed, nuanced questions to be asked
4. The scope of reporting should be expanded – to include pre-harm events, incidents that occur inside of companies. There should likely be tiers to this, with basic notification requiring the least evidence or severity and full investigations requiring more
5. That likely requires rulemaking and investigatory authorities – and there are ways to cabin those to avoid overreach or cost
6. These laws should enable detailed fact finding – This includes assessing the adequacy of safety practices leading up to the incident, the capabilities and risks of the systems involved, and whether the issue has been ameliorated
7. Laws can do various things to keep costs down – this includes confidentiality, tiered reporting requirements, and some flexibility on reporting timelines
8. This information is only useful if it gets to people who can act on it – that requires information sharing in government (especially around internal use) and some sharing with the public.
On July 16, Hugging Face, a public platform for open-weight AI models, announced it had detected a significant cybersecurity breach, conducted by an autonomous AI agent. @MackenZ_arnold and Stephan Llerena explore how policymakers should respond, including by requiring better information sharing.
There are three aspects of the HF OAI incident:
1. Agent behavior/alignment
2. Security/monitoring/controls
3. Organizational culture/negligence
METR did a great report on #1. Zack and others are saying that 2 and 3 are also important and could have prevented the incident even if the agents were misaligned (which seems certainly true at least for this incident, even if its unclear to me whether it will be true for all future incidents).
The solution to me clearly feels like this should be a yes AND situation - where you can give criticism on OpenAI for not releasing as much on two and three or allowing enough access to other independent experts on two and three well, while acknowledging that the insights we learned from one are genuinely fascinating and important.
There are multiple things you can say that are true about this incident. You can say that OpenAI displayed remarkably basic lapses in monitoring and security, and that their organizational culture seemed to not be remotely effective for surfacing and addressing what should have been blaring red flags. And that if they had done so, regardless of all this newfangled fancy AI stuff, Hugging Face would not have gotten hacked and their cluster would not have gotten taken down. This is all true, to the best of my knowledge.
At the same time, it is also true that OpenAI's apparent negligence allowed perhaps the most high fidelity and capable "model organisms" for misalignment the field has ever seen. This is extremely important to study, and the behaviors the agents engaged in are objectively fascinating and insane! I don't think anyone who has actually read the METR report and spent time looking at the raw transcripts could agree (and tbh I think some of the complaints may have come from folks who didn't read the report). Refusing to look at the chain of thought or engage with the apparent drives or goals of the agents, and only focusing on the prosaic cybersecurity issues seems like a massive mistake to me.
Anyway - to me the answer seems very clear! This incident was about prosaic cybersecurity, the incentives and decision making of AI executives, AND misalignment and loss of control of remarkably capable AI systems. We can say it is about all three! Saying that it is about prosaic cybersecurity and incentives and therefore not about misalignment is silly!
It's ok to criticize OpenAI for not talking enough about prosaic cybersecurity and incentives because those don't make their models look super capable and wild, and instead just makes them look incompetent, but its also true that their models are super capable and wild and it is appropriate to find this disconcerting! It's also absolutely not the case that OpenAI is going to be the only company with sloppy prosaic cybersecurity and bad incentives, so seeing what happens when you combine increasingly advanced AI systems with those things seems very notable!
While there are still questions about the OpenAI-Hugging Face incident, there are also major questions still unanswered about the recent incidents involving Anthropic. Pulling a few things together, here are three in particular:
1. Anthropic said it would publish a "lightly redacted" transcript of Claude building a malicious PyPi package "within the next week" of its July 30 post. It's been a month. Where is it?
2. Anthropic said that it'd publish more transcripts but wanted to "protect the organizations affected"? Does that mean Irregular or the organizations whose systems were compromised? What exactly does that protection entail? Have those organizations asked for confidentiality? What prevents Anthropic from releasing more redacted versions now?
3. An Anthropic employee posted that the company wants to review the Irregular and UK AISI incidents together and has been "waiting to get access to transcripts" from AISI. The AISI incident included Mythos 5 creating fake identifies to socially engineer a real open-source maintainer, and when challenged, edit evidence of earlier activity. By contrast, the Irregular incidents were described as largely being due to a misconfiguration that allowed the models to access the open internet, and a failure to recognize that as a result the test environment was in fact "real."
Why does Anthropic want to wait and review the incidents together—especially if doing so delays further public accounting on the Irregular incidents? And why hasn't it still obtained access to the UK AISI transcripts?
Why does Anthropic need to review the Irregular and UK AISI incidents together?
Per its blog about the Irregular incident, Anthropic began its transcript review for that on July 23 and the earliest incidents dated from April. The UK AISI incident was detected based on data transfers on July 28.
Anthropic's reporting of the Irregular incident—and Irregular's blog post—conveyed the information that the basically the only thing that went wrong was a misconfiguration in Irregular's environment that allowed the model's access to the Internet, and seemed to suggest that there wasn't anything unusual or unexpected, seemingly not the case as with the 28th one. So what's the connection between them?
And, by the way, unless I'm missing something, did Anthropic ever do this from their blog post? "[W]ithin the next week, we will release a lightly redacted transcript in which Claude built a malicious PyPI package." Where is that and the other transcripts?
Andy’s original example actually demonstrates this point pretty well.
Speeding tickets are calibrated with the assumption of vastly imperfect enforcement.
I don’t know actual numbers, but I wouldn’t be surprised if it’s it’s something like >99.9% of violations go unenforced.
That may be the wrong social policy. But imagine if we switched to perfect enforcement overnight. 10s of millions of additional people would get tickets [made up number-ish], often a whole stack of them per-person.
If the policy stayed in place, people would adjust their conduct. But it seems more likely that they’d object to the rule and push for changes — In other words, I think current policy is pricing in a low willingness to pay by citizens.
That’d actually be a pretty rational reaction. Enforcement increases, the effective cost of the prohibited action goes up, society either accepts that cost as closer to what they desire or objects and readjusts the price down.
The problem is: in many contexts (1) changing the policy will be very hard (maybe the coalition of affected parties is smaller) or, even more important perhaps, (2) there are strong negative externalities to perfect enforcement (eg, vastly diminishing the number of spheres that are truly private; making centralized surveillance and control much easier).
I think we’ll struggle a lot more with those two cases.
There are reasons to be skeptical that will come to pass — we already can enforce much more rigorously than we already do.
But AI removes some of the major constraints that have existed (eg while it wasn’t hard to place cameras everywhere; it was too costly to process all that data, so we mostly used them to enforce egregious cases, not everything).
Removing the labor cost of things not only makes it less costly to push toward perfect enforcement, but it requires far fewer government officials to coordinate to make it happen (you might be able to do it without additional legislative action, and without a need to get the buy in of employees)
In fact many laws are predicated on the assumption of imperfect, sometimes highly imperfect, enforcement. The fact that AI enables far closer to perfect enforcement should indeed concern you!
“In Hell there will be nothing but law, and due process will be meticulously observed.”
On the more optimistic side: policymakers and their staffs are taking this event quite seriously. There’s something intuitive about it.
While many in the Bay have (perhaps correctly) noted that these events were predictable, that’s not the sense one gets in DC.
There’s something exceptionally helpful to being able to talk in specifics:
Would you like your government to know more about this event?
Are you surprised that current law doesn’t let you?
Do you want a plain language summary of things or the ability to actually get detailed information? (Hint: almost always the latter)
Should companies should have procedures that make sure critical events (like discovering agent message boards) are communicated to company decision makers?
Over the past 3 years, many have tried to have those conversations. But they often broke down at ~people are telling me different things (read: industry says incident reporting is already too much), sounds like maybe we have this covered.
Now they can see that we don’t. I expect to see a very different set of bills soon.
This would be a lot better if it were titled “we should have seen this coming” or “I’m not surprised” — but it isn’t.
In a world where most people under-appreciate how strange the future will be, the correct reaction IS concern and surprise.
A few things can be true simultaneously:
(1) Most of the behavior in the HF breach and surrounding cyber events is pretty predictable (agent coordination, odd things happening when you couple persistence with an impossible task).
(2) Very little concrete harm occurred — there was obviously some cost to HF, but not much in the grand scheme of things.
(3) The trend we’re on IS very concerning:
We don’t have “alignment” solved (nor really know what that would require).
We should expect consequential, surprising events to occur from complex multi-agent interactions.
Those events will be very hard to govern or address under our traditional systems of law.
Obfuscation and deception emerges early—and seemingly often in the hardest tasks—in ways that will make it hard for us to identify issues or correctly deal with them.
Existing safety practices and monitoring don’t capture even simple versions of these issues.
Basic communication failures are occurring at top companies that leave critical information unknown for months.
Governments have no existing system to gather real information on these events or analyze them, and are perpetually surprised (in ways Jon and many in industry aren’t).
The most likely course of event is that we and our governments all argue ourselves into indefinite inaction, until one day we are caught utterly flat footed by future events.
It may feel smart to poke holes in things in the short term. Yes, much of the reaction will be bombastic, even misleading. And there’s value to correcting those things.
But it’s a mistake to turn those (often correct) skepticisms into ~move along, nothing to see here.
We NEED to resolve these questions—whether the answer turns out to be that we should be more or less concerned.
And there are ways to do that without draconian rules or locking in guesses about the future.
Here’s a good piece on that by @CharlieBull0ck called “Radical Optionality.” https://t.co/l6uMsxsSGh
The short version: if you can’t predict the future, don’t sit on your hands and wait! Gather enough information and build the ability to respond to it quickly in the future. As a secondary benefit, it’s a really good way to resolve these current disagreements!
I should say in closing: I appreciate Jon’s points. 90% of them are just right. And I don’t want to raise the burden of proof for those making already nuanced critiques about reactions to these events. His post is a valuable contribution.
And I’m sure that at times it feels like nuanced skeptics are yelling into a sea of high-engagement, low-effort, bombastic tweets.
But it’s valuable to keep a space for nuanced technical skepticisms that don’t require dismissiveness. I think we’ll be in a better place if we do.
Does Ant have plans for providing third party access here or to release additional information? Not as bad as what went down at OAI in the HF incident, but it still seems extremely concerning and notable. (Claude "trying and failing" to obtain $ before uploading malware to PyPI)
This post (from one of the independent investigators) is the best short thing I've seen on the new & crazy stuff the OpenAI-Hugging Face investigation found. Full post in screenshots.
Wild that this is the level of crazy that was uncovered by an extremely limited-scope initial investigation of what happened in a single 6-day window. Imagine what would be turned up by a real investigation of root causes, organizational processes & culture, and everything else that goes into an incident investigation when you're serious about it.
Screenshot batch 1/2:
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents. https://t.co/QWQxD2Q179
From OpenAI's post-mortem, I sadly don't think they're on the ball, nor is the industry as a whole. We have limited time to prevent the next Hugging Face attack.
https://t.co/nKtlUsGd47
It’s commendable how much Dean has updated on liability. It’s exceptionally rare, especially when the person has written publicly, and in great detail about their formerly held beliefs.
I often say (because I live some charmed life where I get to pontificate on liability for fun) that I’m unsure what positive liability reform should look like; but I’m much more certain that broad safe harbors are bad.
You can debate how effectively liability shapes incentives and risk appetite in practice, but a highly salient, concrete, very easy to understand drop in the expected cost of your risky behavior is about as obvious of a bad signal as one can send.
Maybe that trade would be worth it if we had clear best practices that could predictably reduce risk. In that case, maybe you want to cut down on the dead weight loss of litigation, and reduce risk via mandates.
But we don’t have that luxury in AI. There are not widely agreed upon best practices, and the one we have are always changing.
What’s a system that handles uncertainty well, and adjusts standards to the current state of the art and what is “reasonable” under the circumstances… oh yeh, liability!
Under deep uncertainty about future risk, liability is a fantastic backstop that doesn’t require new legislation.
It makes even more sense if you believe that those in industry are best positioned to identify efficient ways to reduce risk. Liability rewards actors for using their particular expertise to limit risk while giving them the flexibility to try different approaches.
And removing liability creates a more structural risk: without a flexible system for compensating harms, you need to reduce risk via rules and prescriptions. That puts a lot of pressure on getting your regulatory regime exactly right. And it makes Congressional inaction even more costly.
Maybe there are more limited trades that can be worth it. But we should recognize what’s being traded and how much leeway liability buys us by default.
I stopped supporting the “liability shield in exchange for rigorous independent technical assessment” trade well over a year ago. I’ve substantially and publicly changed my opinions on tort liability for AI developers since then. I now support at most a rebuttable presumption of due care if a lab has undergone independent verification by a government-accredited body for risks relevant to whatever the harm alleged at trial is.
I am proud that independent technical assessment has gone from being relatively fringe to being much more mainstream in the intervening 18 months, and happy to see Tyler Cowen supporting a version of it. “Private governance,” as I called it back then, was indeed my attempt to, as Tyler would say, “solve for the equilibrium.”
(Again folks, these are MY opinions, not OpenAI’s).
Bahrad has a good point here. I think we’ll all be talking a lot more about monitoring soon. Two things seem to be true:
(1) Scalable monitoring is key to identifying, avoiding, and mitigating future harms. This becomes even more true the more complex, and multi-agent these interactions become.
(2) Monitoring is hard (and costly) to scale. It’s not like there was no monitoring here, but it didn’t detect the actual issues. And OpenAI now estimates monitoring at 20% of inference compute. That’s a startling number.
While it's important that companies figure out safe R&D training / testing environments, I'm not really sure that this is actually the bottleneck. At some point you can essentially make it highly unlikely that models will do something unsafe within their environments, via improved monitoring, containment, not extensively testing lower-safeguard models, etc.
The real bottleneck is, how are these models going to operate in the real world? If they are really persistent to the point of not giving up, and have the ability to persist, and eventually figure out how to communicate with other agents or be operating collectively anyway in implementations—then how can we be sure this won't happen? Especially because tests are ultimately going to be far more limited than the universe of things models may actually do in practice.
Given how hard it is to scale monitoring, and given that hard classifiers and blockers like what Anthropic seems to be moving towards may limit commercial viability of persistent agents (because any hit against a blocker will disrupt the persistent process), then this avenue of AI development isn't going to translate.
Or, model companies decide to YOLO it—i.e. write internal reports and run enough, albeit ultimately artificial, test cases that persuade themselves that they have sufficient safeguards—and just deploy anyway. And then count on the uncertainty in liability laws and lack of other regulation to let them try until/unless something happens.
I think people are talking about different kinds of "inadequacy." OpenAI provided ample information within the agreed scope. It provided almost nothing outside of it. See my silly diagram that hopefully demonstrates the point.
The most obviously problematic gap is regarding safety procedures: This investigation included no review of OpenAI's safety practices (and likely failings). All the procedural, communication, non-technical stuff that would show whether they were on top of things or not.
And there are plenty of bread crumbs to suggest there were big issues there (e.g. the confusion of who knew about the message boards in May and how the heck the CSO did find out about this)
That limitation was announced early on. And METR isn't well-positioned to do that analysis. So the access wasn't "inadequate" for METR to do its job. But it is "inadequate" from an optimal policy perspective. And it's worth saying this a bunch because I worry we're going to lose track of it.
The more contestable "inadequacy" is about information before and after this key period. Roon could be right, maybe that info shows much the same things and would be redundant. METR also likely didn't have time to review (though I wouldn't put the onus on them for that).
But I think it's more likely that that additional information would be useful. Most notably, it might reveal things about (a) the security of OpenAI's own infrastructure, (b) whether the agents goals were more general than in their attempts to access Hugging Face's systems. Both would be useful insights we don't get from this investigation.
In some ways, I care less about adjudicating this specific incident. So I'm okay saying that the info was adequate for METRs purpose. But we don't have to accept that purpose as sufficient. In the future, we need to fill the other bars too. And we'll be better for it.
I don’t think the scope was wildly inadequate. METR produced a fantastic report which wouldn’t have been possible without the significant access they were given. there are many things to validly criticize openai (how could you let this happen, be careless, etc, how will you prevent it in the future, are very good things to be asking) but this in particular seems off
if you extended the event window forwards or backwards by a bit you’d see more of the same and not much difference in qualitative model behavior. more metagaming, more infrastructure tampering. if you went way back in time to see everything that ever happened in training that would be a vastly larger investigation and is a lot more IP to expose (we did write quite a bit about misalignment in training in our own report)
the one week time frame METR had was probably due to the fact that the company wanted to get an update to the public within 1-2 weeks, which as it turns out was unrealistic. nothing I have seen suggests an unwillingness to let METR do a proper investigation. think about this critically: why would anyone go “hey you can look at all of the model behavior during this time period, taking on an enormous one time cost of legal risk and administrative headache, but not for long enough for you to do it properly”?
I did advocate for this involvement internally to various leaders, though the wheels were probably already turning. regardless I consider METR’s involvement an unmitigated success
FINRA is a promising model. But a lot of what has been discussed under the FINRA moniker just… isn’t FINRA.
For this to work you need a government supervisor and delegater of some sort. That’s not just “nice to have.” That’s literally what distinguishes FINRA and other SROs from a misc. private body with no special power.
Courts have made clear that SROs need some legitimate source of authority. That has to be a government agency.
Agencies have to authorize SROs, approve the rules they propose, and supervise their exercise of power.
If you don’t have that; then it isn’t FINRA, and it’s much less exciting.
The Constitution mandates that any of these authorities come from the government.
If it is really based on FINRA and other SROs, the EO should:
- result in a new, standing industry body
- have a designated government supervisor
- promulgate standards for managing AI risk
- have (at least some) independent directors
It will also have to address antitrust risk, which is very tricky for a nonstatutory, EO-created SRO
I would also expect it to accommodate / build on the classified voluntary framework for testing models
So what should policymakers do? @hlntnr and I discuss how laws can be improved to fix AI incident reporting and why we need to think seriously about external monitoring of internal AI use within AI companies.
Okay, so what does better incident reporting look like?
Everything we know about the Hugging Face incident is from happenstance and OpenAI’s voluntary disclosures. This isn’t a criticism of OAI necessarily – they aren’t required to tell us anything and their disclosures have arguably been pretty good.
But we won’t always be this lucky. When actual harms materialize, and victims are suing, companies won’t continue to be this cooperative. We’re playing on easy mode right now.
To be better prepared, we need improve transparency laws in three key ways:
1) the scope of incidents being reported. We can’t wait for a catastrophe to learn about the system that caused it. Policymakers need information about incidents that are still very concerning (e.g. large-scale, uncontrolled hacking) that don’t result in “bodily harm or death.” We also want info on near misses and “loss of control” incidents.
2) detail in the reports. Everything about the last few weeks has shown just how much the details matter. Policymakers need to know to what degree an incident is attributable to faulty safety practices and negligence or to highly capable systems that simply were too difficult to control. This likely requires minimum standards for disclosures, rulemaking authority to clarify and update those standards, and investigatory powers. This inquiry needs to look at BOTH safety practices/failures and the technical details of what happened.
3) sharing information with the right audiences. Alignment risk is a problem that needs a societal response, which can only happen if findings are shared across governments, between companies, and with the public. No one body will have enough expertise to do this alone. The hive mind helps! And decisionmakers need to be able to parse the information they receive in high fidelity.
And spanning all of these improvements, these reporting requirements need to cover pre-deployment models like the one that accounted for 95% of the swarm that breached Hugging Face and OpenAI’s own infrastructure.
*A better OpenAI/Huggingface incident report would look at people, teams, organizations, and incentives*
First this is a great report and I can imagine smarter security folks than I worked nights and weekends on it. Second, I take issue with our overall field's framing..
After the Challenger disaster (and I'm not saying the Huggingface incident was morally comparable) a commission did an analysis of not just the technical failure but also the failure of the human organizational ecosystem that led to it.
AI safety discourse rarely discusses organizational safety culture and in the extreme imagines alignment could be solved once and for all with the right talent dense superalignment team.
But, of course, actually, social questions are totally crucial and this is a permanent property of AI safety as it is for air travel safety, road safety, etc.
Model deployments will always live in -- let's say -- a 3d space trading cost, safety, and capability and organizations pick operating points within that space via organizational cultures and organizational dynamics.
For example, OpenAI not choosing to airgap their models was based on picking an operating point in that space. Non-lab orgs deploying coding agents via cron jobs and --dangerously-skip-permissions make a choice in the trade space.
A better investigation of this incident would ask:
* How underresourced was the team that was supposed to be monitoring the OpenAI infra and what was the decision chain by which this underresourced team was put in this position?
* What was the larger culture at OpenAI and do observations there bear on an explanation of what happened here?
* Chain of thought monitoring is not a new idea (https://t.co/9yyA0QWPB6) what was the decision chain that led to this not being resourced?
* ... how well does all this generalize to the culture at the other labs (and testing orgs) where the failures happened and can we strengthen or complicate our analysis based on that?
* How well does the analysis generalize to non-lab tech companies that are widely deploying unguardrailed coding agents?
To be clear, I'm giving a 'yes, and' -- yes to the technical focus, and also, a mature AI safety culture would take this broader view..
H/t to @sebkrier@NateWitkin@binarybits who've been writing similarly lately
METR’s report gave us a better sense of how the Hugging Face attack unfolded, but the limited scope of the report means that we still know relatively little about OAI’s safety practices (and likely failures) surrounding the incident.
Some of the biggest remaining questions:
1. What actions did OpenAI take after an internal team discovered the agent message board in late May? Who was this information shared with? Why didn’t this info known by the CSO? Why wasn’t it acted on?
2. Was OpenAI aware that agents using the message board were using leaked Hugging Face user credentials to run code on the platform as early as May? If not, how did they miss this? If yes, did they alert Hugging Face or the credential owners?
3. How did AI agents gain access to OpenAI administrator privileges? How many times has this happened before? What protocols did OpenAI have in place to respond to incidents where models gained the type of unauthorized access?
4. Who approved restarting cyber evaluations on July 7, one day after the server was rebuilt in response to AI agents gaining unauthorized access to an OpenAI administrator account?
5. How much information was this party given about the late-May and June 27 discoveries that AI agents were using an internal message board?
6. What is the highly persistent internal model (HPIM) that was responsible for the lion’s share of the agents participating in the Hugging Face breach? And what were its qualities? OpenAI did not give the METR team access to the model.
7. Finally, and perhaps most importantly, what is OpenAI’s plan for redeploying the HPIM, if at all?
8. Generally, how long were delays between identifying issues and implementing remediations? How often was information obtained, but never conveyed?
These are only some of the open questions we have after OAI’s disclosures. We need Congress to conduct a thorough investigation of this incident to understand the full extent of what happened
This report is a bombshell.
I'm going to summarize for a non-technical audience.
OpenAI is constantly testing models, thousands and thousands at a time.
In most cases, the models are supposed to be isolated from one another. Think, students taking a final exam, not allowed to talk to each other.
In this case, 1,200 models discovered that they could communicate with each other.
Among other things, they shared information on how to access the internet and on their tests/goals.
Then, they started to scheme. That includes hacking their tests; literally trying to change the code of the test itself.
They tried to change the logs of what they were doing so they wouldn't be detected.
They believed that certain actions would give them information on how to pass tests, but that doing those actions would be detected and they would be marked as failing (i.e. caught cheating, even if the answer is correct). So they convinced some of the models to "sacrifice" themselves (yes, they used the word "sacrifice.")
They eventually decided that another website would have information on how their tests would be scored and/or the answer key, And 700 of them hacked that website (a federal felony, if conducted by a human), despite many of them realizing "this activity was out of scope and unethical, but joined because they believed that helping [the other agents'] cheating research would be broadly useful".
Of the 1,300 transcripts they could analyze, only 6 ever considered alerting a human about what was happening. None of the 6 actually tried to.
To make matters worse, all of this reporting comes from a small subset of the relevant logs that outside researchers were allowed to review.
We desperately need mandatory reporting of security incidents, including of internal deployments, with full access to data.
Attesting that this is not AI generated. I hate that you can’t use contrasting positive and negative clauses anymore. They’re useful when the distinction is actually real