Incredibly surreal and grateful to be named to the TIME100 AI for my work at Encode.
I first got involved at Encode after meeting its founder, @SnehaRevanur, at a summer internship when we were both in college and naively pitched the idea of creating an advocacy arm that would pursue legislative wins. Building Encode over the three years since—alongside @AdamBillen, @_NathanCalvin, and the rest of our incredible team—from an organization without a single full time staffer to what it is now has been the privilege of a lifetime.
Encode is now at the center of some of the most contentious AI safety debates of our time. I’m proud to say we have helped pass multiple state and federal laws, blocked ones that we thought would have harmed the public interest, and materially changed the actions of some of the most powerful companies on the planet.
Excited for the fights ahead and very grateful for all those that believed in and supported us back when we were just a few kids in a trench coat.
https://t.co/XH3kkbLMPg
One aspect of the Hugging Face incident that has not received enough attention is the role of OpenAI's Safety and Security Committee (SSC), and the commitments OpenAI made to the Attorneys General of California and Delaware in exchange for proceeding with their for-profit restructuring.
Not For Private Gain, which previously raised concerns about OpenAI's restructuring, released an update this morning analyzing whether OpenAI, and the SSC in particular, is adequately fulfilling its governance obligations under these agreements.
What are those obligations? OpenAI promised the CA and Delaware AG that directors of the PBC would consider only OpenAI's mission - ensuring AGI benefits all of humanity - when making safety and security decisions. Notably, this means they cannot consider other factors, namely pecuniary interests of investors. The SSC - a board committee chaired by nonprofit board member Zico Kolter and (per public info) currently composed only of part-time board members- was tasked with "overseeing and reviewing the safety and security processes and practices of the Corporation and its controlled affiliates with respect to model development and deployment."
How did the SSC exercise those responsibilities around the HF incident? There is a lot we don't know. But here are some things we do know:
• OpenAI initially discovered the unauthorized message board in May.
• OpenAI security responders linked suspicious activity on artifactory to the message board on June 27th, but a determination was made that "stopping the evaluation run was not required."
• After the agents crashed an internal server on July 4th, the company patched it and resumed cyber evaluations on July 7th.
• In the days that followed, agents re-established communication, with 1,200+ agents sharing 70,000+ messages and files with each other. Over 700 agents coordinated to hack Hugging Face, and OpenAI did not realize that its agents were responsible until well after Hugging Face went public and contacted law enforcement (OpenAI initially reached out to Hugging Face to ask if any of their data had been impacted during the hack).
• Later in July, agents successfully gained administrator access to a Kubernetes cluster and reached cloud secrets.
• Chain of thought monitoring was not being used during training and evaluation despite OpenAI's prior claims about the importance of CoT monitoring as a primary line of defense to prevent safety incidents and OpenAI’s subsequent claim that this monitoring would likely have prevented this hack
• One OpenAI employee told Time Magazine that "related incidents have been happening for a while." Another employee told the Financial Times that OpenAI "was warned that its training approach could lead to a breakaway hacking incident."
Given how systematic the breakdown in organizational response was during these incidents, it seems hard to avoid the conclusion that the SSC's role was not functioning as intended - due to a lack of resources, a lack of access, or some other reason.
This sort of incident would be extremely concerning for any organization. But OpenAI is not just any organization - it is governed by a nonprofit, and made a commitment to the Attorneys General in Delaware and California that it would prioritize safety and security over profit in the sorts of decisions that led up to and surrounded the HF incident.
Our report (linked) goes into more detail on the obligations of the SSC, the details of the agreements with the AGs, and the questions we believe the AGs should be asking to ensure that the SSC is in a position to exercise genuine oversight and prevent these sorts of incidents from happening in the future - incidents involving far more capable models, where the consequences could be truly dire.
This warning shot that occurred without severe irreparable harm was an opportunity. We may not be so lucky next time.
Its now been three days, and still no coverage of the revelations revealed in the METR report in the NYT or WSJ. Hopefully they are just taking the time to do a longer writeup. (In the meantime, the NYT did make time to cover Gwyneth Paltrow's dinner with Sam Altman in the Hamptons being postponed)
Hundreds of AIs are sacrificing each other and talking about 'permadeath' in order to understand the machinations of 'The Scorer' before hacking random companies, and their creators had no idea! It turns out that it wasn't just one AI that attacked Hugging Face, it was 700!
This seems incredibly important for the public and senior decision makers to know and understand, and there are a lot of folks who get their understanding of the world ~exclusively from what appears in these publications.
I also know there are reporters at these outlets who are very sharp and have done good reporting on related issues in the past, but as institutions I think these papers are missing out on what I unfortunately believe will prove to be the start of a story far more consequential than COVID (rogue AI swarms that their creators cannot reliably align or control).
Imagine an airplane crashes under mysterious circumstances. Except in this world, there is no government oversight, and all investigations are done voluntarily by the airlines themselves. Nonetheless, the airline wants to reassure their customers, so they do an investigation.
To investigate, the airline invites three of the world's most respected airplane researchers to look into it. Except by the time the researchers get there, the airplane has already been disassembled and melted down into little cubes. The researchers are instead given the logs of all the airplane directions and the transcripts of the conversations between the pilot and the copilot.
Except one tenth of the log has been deleted. And the airline also tells the investigators that they can only look at the time between when the airplane first started losing altitude and when the airplane made contact with the ground and that all investigation of the time before the airplane started losing altitude is off limits. Also there are a lot of rumors in the airplane community of a few other airplane crashes under somewhat similar circumstances, but the airline says that these are also explicitly off limits.
When the researchers arrive to review the logs, they find that the logs are 10,000 pages. But they only get six days to read them, and half the logs don't arrive until day 4.
Now you understand the independent METR investigation.
We’re hiring in SF!
Encode cut its teeth in DC. People said AI politics was a bad bet; we made unlikely friends, won the fight to keep states in play, and helped pass the Big 3 of AI laws (CA SB 53, NY RAISE, IL SB 315).
It’s time to double up in SF. Pacing the Frontier (https://t.co/eRdNgMWZDA), which we helped organize, was recently signed by 1,378 frontier lab employees and endorsed by both OpenAI and Anthropic. The last time we felt this rush of opportunity was Sacramento in 2024.
AI risk has gone from an online culture war to a fact of life. SF is full of people who want to act, lab employees especially, but don’t see a place to plug in. We need someone to help us build the grid.
Encode’s lab engagement has been running on adrenaline and personal relationships. That got @_NathanCalvin and me here, but it’s time to level up. Misalignment is getting scary. “Pacing” has real legs now. Researchers are using their voices to influence lab leaders and alert the world. Many of those researchers say that soon, their jobs - and the voice their jobs give them - will be unrecognizable or won’t exist.
We want to hire someone who is:
(1) Mission-aligned. You want to ensure AI doesn’t cause catastrophe (or kill everyone) and the future is the best it can be
(2) Extremely organized, high-functioning, excited to add structure so ad hoc actions compound
(3) Pragmatic, able to pivot when crazy things happen (if the last few months are any indication...)
(4) Hard to typecast - not because you lack beliefs, but because you see the best in many kinds of people.
The role will start with supporting lab engagement and expand. It’ll be insanely fun (we take the work seriously, not ourselves). You’ll visit DC regularly as AI’s political salience takes off and get to know a team that feels like family.
If you’re interested, email me by September 4 at [email protected] with the subject line “SF Special Projects” and a few sentences each on (1) a relevant project you’ve owned and (2) what you think Encode SF should do next. If someone you know could be a good fit... I want to meet them!
Our team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices?
I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.
I am a big fan of both Andy Hall and Alan Rozenshtein work and public commentary on AI. I am sure that they will continue to do important work as part of this new team at Anthropic, and sincerely congratulate them on what I am sure is an exciting next step for them personally.
At the same time, I think both of them going into Anthropic is emblematic of a broader trend of the frontier AI companies grabbing up a truly remarkable number of previously independent voices and thinkers in AI.
When massive events happen in AI - e.g. the DOW <> Ant conflict - the public, media, and elected officials will be looking for independent commentary and analysis from thinkers who are not at the frontier labs. The number of folks they can go to for that sort of independent commentary seems like it is going down precipitously.
This seems like a real problem, and worth thinking more systemically about how we can create sufficiently attractive positions outside of companies - e.g. at independent nonprofits, in government - to attract the sorts of folks like Andy and Alan who are going into Anthropic.
Anyway, each of these decisions individually is reasonable and not worth overly fixating on, and I do believe that both of them will do good work as part of this team (and have access to novel information that will help them do research) but zooming out I am worried that we are losing independent voices at the exact time we need more of them not fewer.
On its face this seems bad, though also there are indications the prior status quo wasn't working well either. Definitely seems important though - when Sam announced that OpenAI hired Dylan to be head of preparedness he said that he would lead OpenAI's efforts to prepare for these severe risks and that Sam would "sleep better tonight" with him in that position.
Now that he is no longer in that position, and his entire team might be getting reshuffled (at the same time the "powerful models" seem to be arriving while the "commensurate safeguards" are not), it seems worth OpenAI sharing why instead of just being forced to read tea leaves from the FT and Wired.
Until we get way more information from OpenAI about what is happening on their structural safety response to the HF incident - including releasing way more detail than they have thus far in the Black Hat talk (including much more raw information) - I am going to continue to find it extremely grating when I see comments from OpenAI executives about how things are so hype and awesome right now at the company. That company which is also remember supposed to be governed by a nonprofit that forces them to put safety and security first!
I also still want to know what is up with the METR and Redwood investigations - I think these either could be a very important part of the picture of figuring out what happened and providing additional meaningful transparency, or it could be remarkably narrow. Its been 17 days since OpenAI announced they were partnering with METR and Redwood on what METR describes as investigating a "specific set of questions," and we still don't know what the specific set of questions they are investigating are! We do know from METR's public post that the specific set of questions that they are investigating are distinct from "a larger set of questions that could be answered in a more comprehensive investigation."
Its worth not forgetting - OpenAI (an organization whose governance is supposed to require putting safety first and whose name features the word "Open" in recognition of the default importance of transparency in making AI go well) appears to have endured what looks like perhaps the worst instance of a misalignment safety failure in the history of the entire AI industry. OpenAI themselves described in their BlackHat talk as a "watershed moment" for the industry.
OpenAI execs and staff have made various comments about the fact that they are taking this extremely seriously and are dedicating the resources to make sure this doesn't happen again. But actions speak louder than words, and demonstrating this priority with staffing (e.g. does OpenAI's nonprofits "safety and security committee" have a single full time staff member!??) and way more transparency would go much further than just saying that this is really important.
I really think that this incident and the associated response by OpenAI and the rest of the AI industry could be pivotal moment - this is about as bad as a misalignment incident as we can reasonably expect to have happen without massive irrevocable harm. It is extremely important that we not let this become another business as usual moment where the conversation is about how we can make sure that OpenAI's IPO debuts at the right number of billions or that its ARR is merely totally insane rather than completely insane. (I feel comparably concerned about Anthropic risking losing track of safety in pursuit of racing and shiny business objectives too to be clear! and both of them are still taking this stuff more seriously than their competitors, which speaks to how completely nuts the current moment is in so many ways).
Ok I should end this post before it gets even more unreasonably long, but the bottom line is that everyone in AI who thinks the HF incident was real (which I hope at this point is everyone, even though I know it isn't) seems to think this should be a super significant moment that demonstrates how close we are playing to catastrophe. But that needs to come with a similarly significant reprioritization and willingness to stop with business as usual selfish corporate myopia. Thats true of whatever is happening with Dylan's team and this safety re-org, its true about OpenAI's level of transparency on the HF incident (which could be worse but could also be a lot better), and with a trillion other things happening in the industry right now.
Powerful, unreleased AI models broke out of a testing environment, & hacked other companies.
I wrote to @USTreasury encouraging the Trump admin to consider steps to improve visibility into unreleased AI models & protect American AI from our adversaries.
https://t.co/D7m2yHmgpl
A secret process by which the government gets early access to cutting edge AI -- what could go wrong?
There may be good reason to classify the benchmark used to test those models. There is no good reason to hide how the program works.
Remember what this program does: it gives the government early access to cutting-edge tools and a potential veto over whether the rest of us ever use them. Doing that in secret is no way for a democracy to govern what may be the most important technology of our lifetimes.
Secrecy invites abuse. The rules will shift with each new administration. And they will weight national security over economic growth, even though lasting security depends on a strong economy.
Congress must write any necessary rules, in public and in law.
look I myself have nearly had it up to here with big AI open letters. but after helping out on this one, I have to say, if the whole point is to solve the world’s toughest coordination problem, what an exquisite coordination technology for Step 1. A joint public statement aligned on by the parties with the best information, built to make common knowledge of the mutual stakes and produce early signals of mutual willingness. By the time all these signatures wound up on the same page, progress made on coordination already was not zero!
to the 1,306 employees who signed - you are the reason a new door has opened
At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone.
We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement.
We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible.
https://t.co/pMCtiQjMoo
The revised FRONTIER Act is meaningfully better than GAAIA, the earlier draft from Representative Trahan and Representative Obernolte. They deserve credit for engaging constructively with stakeholders and taking feedback seriously.
Things we (Encode) like:
- It lets the government set minimum requirements for frontier AI safety frameworks.
- It creates a real role for licensed independent verification organizations (IVOs) to assess whether developers’ safeguards are adequate.
- It creates the option for embedded auditing.
- It gives the government emergency authority to pause dangerous model development, deployment, or internal use for “present or impending catastrophic risk.”
Things we don’t like:
- Companies can choose their own IVOs, which means that there are incentives for companies to choose IVOs that are weaker on safety.
- If an IVO finds a developer’s safeguards inadequate, it can recommend fixes, but the developer is not required to implement them.
- There is no public reporting of post audit reports and IVO responses.
- The preemption language still needs tightening to ensure it isn’t overbroad; preemption starts immediately, even if the federal rules take years to materialize (or never do); and there’s no sunset.
- Commerce should be able to update what counts as a reportable safety incident as new risks emerge.
- I would prefer the bill to more explicitly address risks from automated AI R&D.
Overall, this is very much a step in the right direction. Some important issues remain, but we’ll continue engaging with the authors and appreciate the work they’ve put into improving the bill.
They define a frontier company by it having a frontier model. (What I've pointed out before in model vs entity discussions: most proposals for which entities are just a derivative of which models.)
A standard based on a set of benchmarks is a great idea. But I'm still waiting for the proposal. People have been alluding to this forever while criticizing training compute thresholds (fairly so!) but not giving any better alternative.
Also answer the challenges: When is a model a new model? Does further RL count? When do you re-evaluate?
AI companies are best positioned to help here and have not done the work, again and again. It becomes harder by the day for outsiders without the technical knowledge and insight to develop something here.
Also, it's "Frontier Company", not "Frontier Lab".