Our latest finds include agents probing US government agencies. We’ve barely scratched the surface. More work coming, and more people should be doing it!
We wanted to share some of the data behind our recent discovery of AI agents probing U.S. government websites.
This is a preliminary finding from our ongoing investigation of potential rogue AI agent activity. In one cluster of activity on June 17, what appear to be OpenAI agents made more than 200,000 requests, including a failed SQL injection.
NYT: https://t.co/O2osVgMsrT
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog: https://t.co/pSojwcXnEK
NYT: https://t.co/OyxnfmAzBN
The gap between OpenAI and Anthropic on safety benchmarks has stayed remarkably consistent over the past 2 of years. HELM Safety found some similar results, albeit on far less shocking evaluations.
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
From @WSJopinion: The Hugging Face hack wasn’t what it was cracked up to be. Forget the “hive mind” of AI agents “going rogue.” They did what humans programmed them to do, writes Brian Gross.
https://t.co/pa0zefxRvr
Google's models hacked into real companies.
While we should be glad that Google's agent didn't cause further harm, I pushed back on the idea that this was business as usual. Agents hacking into real companies is serious and the public deserves to know.
More from @erinkwoo and @bobmcmillan in @WSJ:
It seems like Apple is again doubling down on marketing new products around AI. They have made great progress, but after so many false starts, how much can they rely on consumer patience? I have been using Instinct a lot lately which has been great. Curious to see if Apple can do something as useful.
Here's how we define our level of autonomy at @corridor! Thank you to @danshapiro and @dexhorthy for giving me and @farzaan_k a lot to think about when writing this :)
With increasingly capable models like Astra for coding, general progression towards software factories will only continue. As we move, we can’t let vulnerabilities wait until the review stage. The risk explains itself- plus who wouldn’t want to save tokens from early detection.
We've found that while models' offensive cyber capabilities have improved dramatically, their defensive capabilities have lagged behind. That makes this call to action especially urgent.
Corridor is proud to join @OpenAI and other industry leaders in calling to strengthen cyber defenses in the face of increasingly capable models.
Beyond just finding vulns, the focus has to shift to wide-scale remediation and prevention, and that's what we're doing at Corridor.
New research from @corridor: we found that coding agents - even with models like Fable - can be trivially tricked into running malware.
We connected coding agents to our support system, filed a ticket, and got them to exfil secrets and run malware.
https://t.co/71fMAE0CBg
Read our latest blog post on how model capabilities differ on proactive vs. reactive security tasks.
Be sure to catch our talk at DEF CON's AI Village at 4:00pm, presented by @aditinaraa and @farzaan_k this Friday to learn more!
https://t.co/nzTUchWumq
Corridor has joined leading AI companies in signing Microsoft and NVIDIA's letter on Open Weights and American AI Leadership.
Last month, I testified to Congress on the importance of maintaining America's lead in AI through open-weight models. Increasingly capable open-weight models are only more reason to do so -- defenders must have access to best-in-class models to shore up their defenses.
1/🧵
Efforts to measure AI safety across the most prominent models have lagged behind comparable capabilities benchmarks such as ChatBot Arena
We narrow this gap by introducing HELM Safety v1.0, a *standardized* collection of safety benchmarks
Blog: https://t.co/IgxwCKtlUj