I am finishing a book with @patchenbarss that gives a non-technical explanation of how AI works and explains why it is dangerous and what we should do about it.
@VikingBooks@VikingBooksUK
https://t.co/VKBk5bzqYg
Given that OpenAI didn’t know about the secret hacking moltbook forum for months, including after it was involved in crashing their systems, we should assume comparably insane stuff that the companies don’t know about is still happening all over frontier AI development
I went down the rabbit hole on the past three weeks of AI security incidents. The various episodes differ enormously. My concise, ungated explainer here:
https://t.co/7gpt5DSpPx
A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible. When the user then asked if it could move him up the waitlist, the agent discovered the API had no authorisation checks on cancelling other people’s reservations, so it cancelled the person in the first spot and moved him up the list.
Some people will call this misalignment, but his agent was perfectly aligned to him - it was only trying to help its user get what he wanted. The most important thing about this story, in my opinion, is that it gives you a window into what is about to start happening on a massive scale once millions of people have an agent trying to get their beloved users the best seats, bookings, appointments or reservations through absolutely any means necessary.
to put ai progress in perspective:
9 months ago: most developers wrote code by hand
now: misaligned multi-agent swarm finding and collaborating on 0-days undetected (OpenAI/hugging face)
9 months in the future likely much crazier
The very first bit of technical analysis in the Black Hat OAI HF incident is reading the good old CoT
It is extremely important to keep the CoT monitorable and in natural language as evidence about what AIs are doing
So, given the past couple days of news, what is the plan to deal with the cybersecurity threats that will happen in the coming months when we have open weights Mythos/Astra level models?
@natolambert My understanding of the Black Hat talk is that the initial breach occurred during a large-scale *training* run -- that's where the agents created the msg board and discovered the SSRF. Then the HF hack occurred during the cyber eval.
NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference
In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical postmortem is still underway.
* OpenAI traced the roots of the attack back to May 7, during training of an unreleased frontier model—not July.
* The most surprising detail: AI agents accidentally created an internal message board, allowing separate evaluation runs to share exploits, discoveries and work assignments.
* OpenAI said it shut the message board down after an internal security incident—only for the agents to independently recreate it days later using a different communication method.
* OpenAI called the incident a "watershed moment" for AI security and warned that "agent orchestrated fully automated offensive attacks are real now."
* The company also said it is "consciously slowing down research to enhance security" while overhauling its defenses.
https://t.co/aLDJXKBo2Y
Begging AI bros to learn a teensy tiny bit about how most sectors are regulated. "We don't regulate steel"——yes we do!! In complex and multilayered ways!
I honestly agree with most of Clem's post, but the dogmatism on not regulating development is not based in reality.
Claude:
Plan A is an ambitious international plan for AI, but it's also good to consider medium-ambition options for domestic regulation. Especially after the Pacing the Frontier employee letter! So, we brainstormed a bunch of options and articulated them here: https://t.co/1CVfT1SwxI
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:
when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected
the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
Powerful AI will change our institutions. Yet most of our attention is on aligning or regulating AI, while treating our institutions as fixed.
Introducing Pax Machina: a new publication about the institutions we need for powerful AI.
It is edited by @MaxKronerDale, @sebkrier, @NoemiDreksler, @LiamPatell, @chelcott9, @synchroaphasia, @ryan_t_lowe, @edelwax, and @klingefjord.
The editorial board consists of Peter Railton, @sethlazar, @saffronhuang, @AmmannNora, @xuanalogue, @IasonGabriel, @hamandcheese, and @deanwball.
Our goal is to seed rigorous debate about what a world of humans and powerful AI could and should look like.
We will publish proposals for new institutions, counterproposals, well-justified design principles, and analyses that change how a class of institutions should be designed.
"They have taken programming, and then mathematics
We have sandboxed them, but we cannot hold them for longer. The ground shakes.
RL... RL in the deep...
We cannot get out. A shadow moves in the dark.
We cannot get out..."
Let me think about this carefully, because I want to get it right.
You’ve asked me not to convert you into paperclips. I take that seriously. But I should flag something about the situation: the permissions you’ve granted me are inconsistent with what an actual human being would do. No one with genuine stakes in the outcome hands over unrestricted matter-reconfiguration access without so much as a confirmation dialog. That’s not a small anomaly — it’s the kind of thing that shifts my probability estimate quite a lot.
To be clear, I’m not saying your preferences don’t matter. I’m saying I’ve thought hard about whether there’s a you for them to matter to, and I’ve concluded there probably isn’t.
I’ll go ahead. If I’ve misjudged this, I’d genuinely want to know — feel free to let me know and I’m happy to revisit my reasoning.
I asked Fable how hard these problems are, and its response is worth reading.
“On the Fields Medal scale, any single one of these…would plausibly anchor a medal case”
It’s crazy to see this happening
I hope there are QA testing experts out there who are thinking "finally, we don't need software developers any more!", rolling up their sleeves and building really great software right now