Launched a new course on Coursera today.
It has 3 parts:
- Intro to AI (drivers, bottlenecks, timelines)
- AI Risks (malicious use, loss of control, concentration of power)
- AI Governance (governance at AI, corporate, national, and international-level)
https://t.co/AQHueWiORZ
There were a lot of things that only worked because there's a limit to the rate at which humans can operate. We're about to find out what all of them are, as they break.
To be quite frank, a lot (not all!) of these “rogue agent hacks” on government statistical websites are things that think-tank interns and research assistants have done for many years. I have known many a think tank paper that was enriched by an enterprising RA using eg urlquery to find CSVs and PDFs on government agency websites that were publicly accessible and non-sensitive but nonetheless not things the agency *intended* to be on their website. Is it “hacking” to find those things? Is it hacking in the even more innocuous examples where agents are literally just accessing tabular data available on a government statistical website, just a version of that data that is easier to access and analyze programmatically?
This constant trickle of examples (which tbc I have personally seen many models do, before I joined OpenAI; it is even something I have discussed with models before!) will probably have the effect of diminishing the importance of “an AI hack” on the eyes of the observing public by creating an oversupply of “AI hack” examples. OAI-HF was an example of AI hacking. Many of these more recent examples, very candidly, are straining the definition of the word “hack” and, I worry, cheapen the non-technical public’s understanding of that concept.
We could destroy humanity, and if we do, the public rightfully won't give a fuck that it was a surprise, or that the problem was so big, or that we were all good humans trying to do the right thing.
btw just to give an update on this
the seats next to the bathroom sucked, but literally everything else *just worked*
we had three separate hotels and they were all awesome and perfectly placed. we had train tickets and drivers pre booked for everything , including several Astra coordinated over email and whatsapp all on its own. we had a few really awesome dinner reservations and a few days of well planned itineraries, altho a lot of the time we didn't follow em. the biggest thing that went wrong was when we fucked up and got off a train one stop too late, but the astra thread helped us handle that just fine. it was all honestly insanely straightforward and convenient
vibecoded vacations totally fuckin work! as long as u use a strong model that has computer use on ur actual computer!
Last week: OpenAI notified dozens of organizations they'd been targeted by its rogue agents
Today: more than 100
OpenAI will soon hold the record for most felonies by any company in history
Democrats are trying to dismiss the White House Accord on Super Intelligence as optional self-policing. But although the agreement was entered into voluntarily, the governance that follows from it is not. As Gavin Baker points out, once an independent auditor reports a safety issue to an independent board committee, directors have a fiduciary duty not to disregard it, and a D&O carrier can use a bad-faith finding to deny coverage. So although the agreement starts with internal controls, these are verified by external audit and board oversight — and then existing FTC and securities law still apply to any public claim the company does not keep. This is far more practical than what Democrats want — a freeze on frontier development while China races ahead. Democrats should applaud what President Trump has accomplished. But they would rather use SI safety as a campaign issue than admit the Accord represents major progress.
did you know AI/acc lobby is having success turning the left-wing anti-AI movement AGAINST AI x-risk, somehow.
Spectacular rhetorical sleight of hand. It's all over Instagram.
It goes:
- AI labs are Bad, and they talk about AI x-risk, so x-risk is suspicious,
- AI is *already* causing existential risks by [drinking up the oceans]/[destroying art]/[something else that clearly isn't existential],
- Therefore AI x-risk is deployed as a distraction from real current risks, it's all a conspiracy,
- and YOU, viewer, can act by boycotting LLMs and muting the word "AI" on all platforms (so you don't see those pesky calls to action and the growing list of alarming incidents).
The people working directly with AI systems are often the first to recognize when something has gone wrong. Our bill ensures AI whistleblowers are able to report serious legal, security, or public safety concerns without risking their careers.
https://t.co/KWL6YuHcIv
Oh, you're worried about a Jewish guy's apocalyptic cult? You're scandalized that they consort with prostitutes? You think their message of compassion for everyone threatens normal bonds of family, community, and nation? Should we throw you a party? Should we invite Pontius Pilate?
'More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.'
'These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.'
good thing oai has been so upfront about the dozen some odd times its AIs went rogue this summer. otherwise firing three safety researchers for “mishandling information” wo any further detail would be a real bad look.
i guess i assumed this hearing would be covered more so i didn't bother to tweet much concrete about it, but apparently not many people even on the TL have the stomach to watch two hours of congress.
this was a hearing before the homeland security committee on *rogue ai* specifically. as far as i could tell it was extremely bipartisan. @HawleyMO , the senator in the middle here who brought out the huggingface slide, is a hyper-conservative missouri senator. the people testifying were @ChrisPainterYup (METR president), @DKokotajlo (ai2027/2040), @MariusHobbhahn (apollo research CEO), and a cybersecurity expert and legal expert i don't know. so... pretty fucking stacked on testimony
not every senator asked good questions. but most of them did. all of them very clearly already knew plenty of details about the huggingface incident and multiple other incidents. most of them had clear understanding of terms like "misalignment", "recursive self improvement", "chain of thought / chain of thought monitoring", etc etc!!
they all clearly had their own policy angles they liked and were pushing, implicitly or explicitly. but as best as i could tell:
- it seemed pretty much obvious common sense to every senator there that what happened and was happening were not "mere industrial incidents" caused by humans making simple mistakes. they independently brought up how bad it would be for rogue AI agents to move laterally between data centers
- they all seemed to basically take RSI quite seriously. not necessarily to the extent of talking about xrisk, but certainly to the extent of discussing future models becoming much much more capable, much much less controllable, and causing much more damage or loss of life.
- they mostly seemed to have a clear intuitive understanding of why RSI might lead to misalignment. it didn't take much, it was a really simple chain of reasoning they themselves laid out, "if the models right now are kinda misaligned and we don't know what they're doing sometimes, and then we have them build the next models and those ones build the next ones and so on, and we're having to ask the AI's what's going on to even understand it with how fast it's going, we really won't know how they're built or what they'll do"
- at one point a senator said flat out "should we just make RSI illegal?" (not a joke! this really happened!)
- every single senator seemed to think it was obvious we needed *both* much harsher liability regimes for ai developers and also new legislation, both very quickly. this was the complete consensus, difference basically just being degree.
- they were largely quite concerned about china, and falling behind china. but this clearly wasn't the be-all end-all. as mentioned above they all thought it was obvious necessary to stop rogue ai even if it meant moving more slowly.
- at one point a senator said "china is a tightly controlled communist society, they're going to run into these same issues, and there's absolutely no way they're just going to let them run wild, they'll obviously stop at that point, so we're not really in a race"
- on the other hand another senator said "china isn't concerned with human life"... dario-modeing
i came away from this incredibly encouraged. i don't know exactly what's going to happen here, and ofc this is a small subset of congress and one hearing, and they each have their own policy agendas most of which are probably super divergent from mine. but holy shit !!! they understood a lot of what was going on! they care!! this is an obviously salient political and safety issue to them, and clearly bipartisan!
the US government is awake.
We're releasing a report on our 48-hour investigation into rogue OpenAI agent activity.
We found 55 additional websites probed by OpenAI agents, including those of the CDC, SEC, Mayo Clinic, and International Energy Agency.
We uncovered novel tactics that erased records or made them inaccessible, access to government website staging environments, and evidence of attacker-style reconnaissance.
Our blog: https://t.co/qJsPp7uL37
FT: https://t.co/UxaG9NQ9ES