I dunno man, it just doesn't sit right with me that there are public technical blog posts that Claude simply can't read or discuss because of "cyber risk". Who does this protect?
*Against today's high friction trusted cyber access programs and against restricting open-weights cyber models*
The key object of any discussion of AI restriction is the tradeoff between attacker use friction and defender use friction. I don't think today's trusted access programs dial this Pareto-optimally. Neither do proposed bans on frontier open-weights cyber models circulating in policy circles.
Take Daybreak and Glasswing from OAI and Anthropic. These programs serve unguardrailed frontier models to trusted participants; I'm for Know Your Customer in principal, but these programs are too restrictive for defenders and buy us too little in attacker restriction for the cost. From public information less than 1% of software developers are in these programs and able to use them to find and fix issues in their code.
It's true these participants cover the most highly leveraged codebases in the world but this still leaves an enormous gap; take the software objects Hacktron stepped through to access OpenAI's monorepo, libheif and the SaaS product Discourse, neither of which, I believe, were targeted by Daybreak/Glasswing. Diffusion of the best and most frictionless cyber capabilities to the long, vulnerable tail matters a lot.
Open weights is also extremely important for defenders. Many organizations can't or don't want to share sensitive security data with OpenAI or Anthropic, often for understandable reasons; from what we can tell from the outside, these organizations seem to be playing fast and loose with security themselves.
And if we're going to make security work at scale, we'll need to distill large frontier cyber models into smaller models that can monitor high volume trillion-event-scale feeds at low latency / high throughput, find and fix bugs cheaply, and monitor security events on-prem, on-device, and for military uses, in the field with patchy Internet connectivity. We need frontier open-weights models for this.
To be clear: I do think we should go to *extremely great lengths* to restrict attacker access to frontier cyber models.
But there are many policies that *are on the efficient frontier in denying attackers and enabling defenders* that we should consider.
Here are things we can/should do to deny attackers AI access:
* Low friction Know Your Customer programs across all inference providers based on self-enforced or regulation-enforced information sharing and blocklists that codify known-attacker signals; restrict attackers while pouring gas on defender adoption.
* More defense in depth; all inference providers should do more to monitor API usage, detect attacker behavior, kick attackers off, and have them pursued, arrested, or sanctioned where appropriate.
* Closed providers like OAI, Anthropic, Google, and Meta should do this, but so should open-model inference providers like Together, Anyscale, Fireworks, and OpenRouter.
* Time-to-detection-and-response re attackers using cloud AI inference should be a KPI any AI inference company shares publicly; at least under a self-regulatory regime but probably eventually under an international regulatory regime.
* It's true that the most sophisticated and capital-endowed attackers can build/access their own inference compute; but these operations should also be detected and shut down as well as possible by monitoring inference hardware markets as well as we can.
* Victims of (AI) cyber attacks should also be induced to share more information so that we can get a moving epidemiological picture of attacker use. Every AI-agent-based attack should also become a public natural experiment for the security community to learn from.
* The poor reporting protocol OAI followed around the OAI/HF incident is not a good sign of maturity here. Why aren't we pushing much harder on public threat-intelligence sharing so defenders can prepare for the AI cyber phase transition?
None of this will prevent nation-states and sophisticated actors from adding agent swarms, backed either by Chinese or American models, to their arsenal of strategic weapons.
But we shouldn't confuse geopolitical problems for technical problems justifying slowing defender adoption and hardening the world's code and infrastructure.
Nation-states will find ways to get access to strategic AI-based cyber weapons with or without Daybreak, Glasswing, and open weights restrictions, because the two sovereign AI powers and their allies are in geopolitical competition. API-access gating is not sufficient defense against state-level cyber threats.
We should absolutely try to slow attackers down. But we have ~10 trillion lines of code to secure with AI as fast as possible, and a ~120trn GDP global economy to apply AI as an intrusion detection substrate over, and high-defender-friction restrictions are a poor stop-gap for accomplishing this...
(34 days out) Excited to share the lineup of the @SANSInstitute AI Cybersecurity Summit agenda and speakers, and why this was built for you:
Every speaker selected to inform and instruct on present day AI security realities.
Panels with our field's best. They are clearly not afraid to have hard conversations about what AI changes for cyber defense.
Workshops designed for Monday morning. And social things so you can say hi to peers and find battle buddies.
Thanks to my cochair Sounil Yu @sounilyu, to our partner event @unpromptedconf, and to our speakers:
Morgan Adamski, Principal, Cyber, Data, and Tech Risk Leader, PwC
Michael Collins, Senior Principal Security Engineer, Mastercard
Sergej Epp @EppSecurity, Cybersecurity executive; creator, ZeroDayClock
Gadi Evron @gadievron, Founder and CEO, Knostic
Jeremiah Grossman @jeremiahg, CEO, Root Evidence
Katie Moussouris @k8em0, Founder and CEO, Luta Security
Amy Herzog, VP/CISO, AWS
Vinh Nguyen, Senior Fellow, Artificial Intelligence, Council on Foreign Relations
John Hultquist @JohnHultquist Chief Analyst, Google Threat Intelligence Group
Daniel Miessler @DanielMiessler, Unsupervised Learning
Marcus Hutchins, Cybersecurity Speaker and Principal Threat Researcher
Ciaran Martin, Professor; Director SANS Cyber Leaders Network
Chris Cochran @chrishvm, Field CISO, VP AI Security, SANS Institute
Joshua Wright @joswr1ght, Director and Senior Security Analyst, Counterhack
Steve de Vera, TDIR Threat Research III, AWS
Kyle Shields, CISM, GCIH, GCPN, Security Specialist Solutions Architect, AWS
Andy Dennis, Head of Field Engineering, XBOW
William Reyor, Lead Solutions Architect, XBOW
Joshua Corman @joshcorman, Executive in Residence for Public Safety and Resilience; Lead, UnDisruptable27, Institute for Security and Technology (IST)
Caleb Evans, Software Engineer, Agents, Airbyte
Trinity Harrison, Principal Software Engineer
Lenny Zeltser @lennyzeltser, Faculty Fellow, SANS Institute
Pedram Amini @pedramamini, Chief Scientist, OPSWAT
Adrian Wood, Security Engineer, Dropbox, BT6
November 2-3, Arlington VA or online: https://t.co/wteGL83uJr
We must find ways to stay relevant and understand the world.
Mine is through innovating in vuln research.
And, of course, if you want to secure your agents - DM me.
Every model has new ways to speak "AI Native". My hobby: Find them - through vulnerability research. That helps me, and @knosticai, stay AI-relevant. I share five below. We live in the most interesting era of history.
The first thing I tried? Convince Claude to find logic vulns
2. Show me how you managed to write a working exploit by using simplistic models
Check out John Cartwright (and maybe me) at [un]prompted on exploitation with Haiku, based on auto-research loops and code understanding
All of this debate over alignment vs sandboxing vs monitoring feels like asking if we need laws, a police force, or just locks to stop theft. We obviously need all of them!