METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine, forensic teams say — first documented case of civilian deaths caused by a Russian drone using fully autonomous targeting https://t.co/CN6f0IOSAI
On July 28, the FCC restricted imports of new foreign advanced robotics over unacceptable supply chain and cybersecurity risks. More is needed: component provenance checks, military cybersecurity testing, and robot model evals.
Read my explainer on IAPS' Attack Surface below.
There's been a decent amount of speculation about whether SB 53's incident reporting requirements cover the Hugging Face breach (i.e., whether OpenAI is legally obligated to report the breach to California's Office of Emergency Services.
Most of the discussion has been around the definition of "critical safety incident" and whether the HF breach qualifies. See e.g.
https://t.co/NggnmywmdU
But I was thinking this over today and realized that even if the HF breach had been a "critical safety incident," SB 53 doesn't require all critical safety incidents to be reported--frontier developers are "encouraged, but not required, to report critical safety incidents pertaining to foundation models that are not frontier models." So if the model(s) that executed the HF breach were trained on less than 10^26 FLOP, OpenAI would have no legal obligation to report.
Seems like something that future incident reporting legislation should fix. Most of SB 53's requirements are developer-based rather than model-based, which is good for a bunch of reasons, but the incident reporting is an unfortunate exception. The bar for "critical safety incident" is already quite high, there's no good reason to exclude critical safety incidents involving sub-10^26 models which could still be extremely capable.
Anyways, it would be good for calibration to know if the model(s) involved in the HF breach were trained on >10^26 or not; hopefully that'll be specified in METR's writeup or in some future disclosure by someone at OpenAI.
This has a fantastic presentation with an easy to understand methodology and implications.
Great fellowship reading materials imo. I will be using it our Policy Fellowship when I talk about distillation and traditional privacy risks.
@ every university organizer, consider including
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
So my main takeaway from the OpenAI Black Hat talk is that today’s “Highly persistent experimental internal-only model” seems perfectly capable of exfiltrating its own weights.
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more.
https://t.co/zUR1dqmuzi
I hope it can answer a lot of the questions folks have, and we will release a full detailed postmortem at a later time!
Frontier models quietly change their behavior depending on who they are talking to.
If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests.
We call this user awareness. 🧵(1/)
This benchmark that plays into LLMs' coding capabilities shows how much greater the risk for misuse is if the actor is willing to narrow down the task. It is, however, still unclear to me if the processing was on-device which would be super interesting to know!
Should we be worried about how good AI is getting at coding autonomous drones?
Introducing Drone-Bench, a benchmark where AI agents code drones to complete a simple autonomous surveillance task.
Drone-Bench is independent but based on Project Pilot, our work with Anthropic.
When thinking about highly generalizable signal-agnostic autonomous AI drones I found strong bottlenecks in the weight, battery expenditure and cost of edge compute because I was assuming some sort of VLA system.
The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. We’re excited to see what scientists and researchers are able to create with our upcoming Astra models!
Was great to be on @cnni to talk about how Anthropic also has rogue AI models escaping that they didn't know about
"You're seeing across openai, across anthropic, across other companies, AI's are just kind of breaking out of these companies left and right [...A]nthropic AIs were also escaping and causing some small amount of harm. And Anthropic didn't notice for over a month."
"It's completely unacceptable to be building a product and then have that product be able to escape and cause harm. It's completely different from other products, like when I'm using a hammer. The hammer doesn't just like go and attack my friends without me operating the hammer in the first place. But these AI systems, they require a whole new level of security.
But right now, we have a whole industry that's just moving so fast. They're constantly trying to make sure their products come out first before their competitors, and they just they don't have any time to slow down and make sure there's good security for their ai systems. And with AI is at the level they are at now, this is just not an acceptable situation."
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
Between this, and the signing of the "pacing the frontier" letter, we're hitting a streak of great collaborations for a world with more secure frontier AI!
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
I am especially confused by the logic behind AI use in this one, since the AI text is concentrated almost exclusively in the policy recs section
- https://t.co/3BfQH0NtjI -
There seems to be a concerning amount of AI-text or humanized AI-text in policy & opinion outputs. I think most people would want policy recs to be human-sourced.
Deferring without disclosure is voluntary disempowerment.