This isn’t how things work in any other incident reporting regime.
Not aviation, not nuclear, not medical devices, not securities.
Even in cyber (where the victims have valid reasons to keep the vulnerability quiet), they get max 90 days.
Even in confidential reporting regimes (like CIRCIA and ASRS), they still release anonymized info.
We’ve worked through these questions before, and the answer is never ~we defer to the wishes of the company.
It is *exceedingly* rare for me to feel any positive emotions --relief, joy, optimism - on this platform.
Reading @DarioAmodei's new essay and the chorus of agreement tweets (@elonmusk@sama) did just that. I read it all sitting next to my 11-year old. Suddenly, the weight on my shoulders felt a bit less heavy.
I love the third party embedded evaluators idea. Bravo to @AnthropicAI for leading on this and @OpenAI for quickly following.
Last weekend, I spent several hours listening to @ajeya_cotra's interviews re: @METR_Evals Hugging Face report. Ajeya came across humble, and likely wouldn't call herself this, but to me she seemed like a superhero. Principled, brilliant, dogged, untiring in her determination to investigate the (wild!) incident and share the findings transparently with the public. Her main roadblock seemed to be the limitations OpenAI placed on her team's investigation (time allowed, scope).
I am heartened by Dario's proposal of many embedded Ajeya's inside the frontier labs -- and with far greater access than what OpenAI recently granted Ajeya's team.
This arrangement rests on a key feature of groups like @METR_Evals: their independence. As is the case across so many challenges AI raises, from democracy to jobs, we need a thriving, independent sector to tackle these issues, outside of commercial pressures.
Candidly, over recent weeks, many of us have felt deflated by the migration of top talent into the labs and the loss it creates for the public square. I hope this moment ignites a public spirited mass movement of rock star talent into METR and other independent organizations like it, which are mission critical to getting this right. And of course into the government, which desperately needs AI expertise that can keep up with the frontier.
I was just reading new @pewresearch data on young people's pessimism on AI and the future. It feels like we need something akin to President Kennedy launching the Peace Corps -- a call for a new generation of talented young people to join this cause to *do* something about their AI concerns and use their talents to safeguard humanity, as Ajeya did, whether via the govt, independent orgs, or in frontier labs' teams working doggedly to get this right.
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators.
the more eyes and people with distributed skill sets the better.
it would be a good idea to fund several efforts on this.
What you are seeing in math right now is a consequence of the jagged frontier, and a precursor of what is to come in other professions. Yes, mathematicians do math, but they also have other tasks they view as important (mentor students, maintain a scientific community, safeguard the future of a field, foster a love of math) that AI can't do.
At least one worry that mathematicians seem to have is that by focusing on the flashiest, most obvious element of what mathematicians do (make proofs), the AI companies are damaging the other tasks that AI can't do. AI can discover superhuman proofs, but that is not all that the math profession is about, and actually can undermine and reduce the attention to the other aspects of the job that are important to mathematicians. It becomes harder to defend the value of the many other tasks mathematicians do to the outside world if the most visible part is taken away.
I suspect we will see more of this across fields and professions that will increasingly be forced to help people understand that their jobs consist not only the most visible tasks that AI can do, but also tasks that the AI cannot do or does badly.
AI as Normal Tech is one of the most thoughtful, coherent pieces of policy writing in the last few years. If you haven’t read it, you should.
Its policy recommendations get far less attention than its predictions though, and that’s warped the messages people take away from the essay.
The authors do not argue that AI as Normal Tech = all will be well, do nothing.
Instead they make the case that policy should focus on “reducing uncertainty” about the trajectory of the technology and its consequences + building resilience.
I think that’s surprise a lot of people who like this general meme. It sure sounds a lot like most frontier AI policy proposals I encounter, and it would be vastly more ambitious than what we have right now.
I’m not saying there’s no gap. But the fact that this essay has been cited to oppose many sensible, uncertainty-reducing policy proposals directly contradicts the conclusions of the authors.
Incident reporting, whistleblower protections, transparency laws — all are cited as prime examples of good policy. And yet, more often than not, I hear this meme invoked to push back against the need for those exact policies.
Did people just miss Section IV of the essay?
No, I think what’s happening is that people want to read a careful, even-handed essay that is painstakingly precise to instead stand in for a general meme that supports their confidently held predictions for the future or to dunk on predictions they find implausible.
But every bit of this essay shows that these authors actually mean what they say — they are carefully reasoners, who want to rigorously evaluate claims, and the same humility that leads them to push back on many claims of impending risk also make them admit their deep uncertainty about the future.
They’re far less concerned about dunking on other people to prove they’re right, and much more focused on how to handle that uncertainty.
So instead of rejecting policy efforts, they argue we should aim for policies that are “robust” to multiple assumptions and help us better understand the future.
This seems empirically, normatively, and strategically compelling. The goal isn’t to bet all your chips on whatever belief you put 55% credence on. It’s to advance policies that handle the portfolio of possible outcomes well and help you adjust that portfolio more intelligently as you learn more.
What this all means is that their predictive claims are less load-bearing than you might think. Sure, they predict AI will be a normal tech, but the practical upshot is that we aren’t sure and there are plenty of policy interventions that are robust to different empirical and normative beliefs. So we should pursue those!
The reality is that most of those policies, which should represent some reasonable, moderate flank, remain unadopted, weak, or highly contested.
And as valuable as arguing on Twitter is, the way to actually resolve these disputes is to build a policy system that allows us to learn and adapt, to “adjudicate it on the facts” so to speak.
If people took that away from the essay, I think we’d be on a much better trajectory than we are now
A few things seem true:
(1) CoT monitorability was always fragile — The relentless tide of tech progress was going to pull us here someday.
(2) It was less clear how quickly we’d get here — effective use of latent reasoning is not an easy task, it took time to get here, and it could have taken longer.
(3) CoT monitorability is actually useful — even with all the caveats. Yes, it’s not ground truth; yes it can be misleading. But just look at METR’s recent investigation — we learned a lot from CoT. Many promising monitoring approaches also rely on CoT. Last week, I saw a few dismissive comments about CoT as akin to trusting what a 5yo tells you. But if you want to monitor for house fires, a 5yo is much better than no one, especially if the 5yo sometimes openly admits s/he lit the fire.
(4) We don’t know how big the change is (yet) — some at OpenAI have already publicly downplayed how much Astra has displaced CoT.
(5) There was a firm boundary, and it’s now ebbing — Less than a year ago OpenAI and others were treating this line as sacrosanct (at least publicly). Coordination is hard, this undermines our it.
We need more ambition in AI safety and governance.
There is more important work to do than there are organizations to do it, and more funding is available for ambitious projects than ever before.
To address the gap, we're hiring Entrepreneurs-in-Residence at @GovAIOrg. Participants get a year of salary plus ~$150k in seed funding via partner funders to start a new AI governance and safety org or project. We take no equity.
Even though it wasn't an explicit goal, GovAI staff have spent their time starting new orgs, including @Safe_AI_Forum and Trajectory Labs.
We're interested in a range of orgs and projects. Examples include: automating AI governance, an organization that rapidly produces well-evidenced, expert-endorsed “proto-standards”, an AI Bellingcat, a sub-frontier model evaluator, a TechCongress for outside the US.
We take rolling applications. The pitch is max 2 pages. If you've built something impressive and have context on the field, consider applying.
@sethlazar couldn't agree more. It has felt at times like the government or the labs are treated like the only relevant actors here, but that construction is far too narrow for an AGI future.
Fear of regulatory capture may be (1) misplaced or (2) exaggerated but there's another threat to the rule of law that warrants more attention and is very much a product of lab behavior:
Epistemic dependence (in brief, the brain drain).
Even if labs are not actively trying to unduly influence legislation, there's still a risk to the rule of law if:
- the government can’t identify the relevant technical questions without the lab’s assistance;
- the government can’t evaluate the company’s methodology;
- no independent third party has sufficient access to challenge the lab’s assessments;
- government officials cannot tell whether competing experts are actually competent; and
- withdrawing the lab’s cooperation would effectively disable governmental oversight.
These are great steps! Here's 8 other things we could do:
1. Congress should fund CAISI at ~$80 million instead of $10 mn, which is our internal analysis of what it'd take for CAISI to actually fulfill the purposes laid out in the AI Action Plan and other Trump admin directives.
2. The NSA, CAISI + others should plan for the moment when >Mythos-class models are distilled or trained in China, and make a real effort in preemptive cyberdefense. We called this last year, and have some ideas on what to do (https://t.co/d4hUMQQZ8p, https://t.co/Vj8gHkBLo8, https://t.co/P3jOcSdFRa)
3. OSTP and NSC should coordinate building RAND-style SL-4/SL-5 security for frontier model weights. Distillation is one way to get somewhat capable models, but stealing model weights gets you the best model, and it's completely doable for well-resourced state-backed actors. The weights themselves are the crown jewels, and most labs aren't close to being able to defend them! Once we train a 10x Mythos soon, we'll wish we had a secure environment to run it in. (More implementation details here: https://t.co/QKESFX7HE5)
4. Relatedly, fund + help staff an insider-threat / counter-intel program for frontier labs. It is much harder to protect model weights if adversarial people have privileged access.
5. The White House should direct Commerce/BIS to strengthen AI chip and SME export controls to adversarial countries, so that even if cyber-capable models are distilled or stolen, they can't be deployed at scale on American chips. China has huge domestic production bottlenecks (https://t.co/LwGx5fvran), so exporting fewer chips makes a difference, pound for pound.
6. And because smuggling is still a problem, we should also be deploying chip security measures like privacy-preserving country-level location verification, which will allow us to export more chips to semi-trusted countries while verifying that they're not being smuggled to adversarial ones (more: https://t.co/iUpSueHCKt), and there is more AI verification work to be done to enable more mutually beneficial trade without national security downsides (https://t.co/Kvx9oKF0vP).
7. On top of funding CAISI, we should direct it to run pre-deployment evals for CBRN and cyber uplift on a classified track. You can't hold adversaries accountable for abusing US models if we don't systematically measure what those models can do in the first place.
8. The NSC, NSA and CAISI should write the emergency-response playbook for the day a Mythos-class weight leak is confirmed, or distillation is successful. Who does what, in what order?
To be in a good place, we should've started years ago. But it'll only be more urgent each passing month.
Compute stock is growing 3.4x/year; LLM inference prices declining at -40x/year for a fixed level of capability; software progress is improving so quickly that the pre-trainig compute we need to reach a capability is 3 times lower each passing year (https://t.co/QGrPUQ3mng)...
These are just some ideas for government, related to distillation and model weight theft. Philanthropy and the private sector have big roles to play as well.
We have so much work to do!
The idea that technology, like AI, deskills us is not a surprise. I learned cursive at school, my father learned how to use a slide rule. Neither skill is widely mourned.
What is important is whether we will make deliberate choices about what skills to keep & which they will be.