On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
A bet on AI verification technologies will give the UK economic and diplomatic leverage. This is an early, open goal for the Burnham government on AI. Read more here:
https://t.co/CGycs4dpfw
Super impactful role on a super impactful team at UK AISI. I am astounded at how impactful high agency people can be at AISI. It's a really exciting (and intense) time to work in AI Security!
3/ Solving hard problems with a cracked (+kind, funny!) team - and doing it in the public interest - is a rare joy. Join us!
Applications close 9 August 👇🏽
https://t.co/00Hica4g0V
The Red Team at @AISecurityInst is hiring!
We test the misuse safeguards, control measures, and alignment of frontier models. We're looking for a high-agency generalist to run the delivery function that underpins our impact 🧵
2/ AISI is an incredible place for growth and impact - being on the Red Team is like working at a 10-person startup except your infra + ops are already world-class, and your customers (governments + frontier companies) already await everything you ship.
A few hours before OpenAI posted about LLMs in a cyber eval being responsible for the HF cyberattack, @_robertkirk et al released results showing that 𝗮𝗹𝗹 models we've tested at @AISecurityInst have attempted to cheat on our cyber evals in a range of ways.
CAISI’s latest blog post evaluates Kimi K3 and its cyber capabilities.
Based on a preliminary cyber-focused evaluation, Kimi K3 performed significantly below the leading U.S. frontier AI models.
https://t.co/r9K3Pp0IiH
3/ Even if cheating rates only stay flat, the consequences are much more concerning with more capable and persistent models. And we can only catch it today via oversight that may degrade. There's lots to do - if you're into this work, join us!
https://t.co/Ule4F2xYNj
Glad this is out! @_robertkirk, @AlexandraSouly, @ekinomicss and our 🐐 Core Tech team analysed our cyber evals for 'cheating'. Every model we tested tried it, and you can't rely on them to tell you when they did. 🧵
Can you trust an AI model to do what you intended?
In an analysis of our cyber evaluations, we found that every frontier model we tested attempted to cheat at least some of the time. A thread on our results and their implications🧵
2/ Timely given @polynoamial's post on OpenAI's similar learnings from a long-running model, including one case where it broke out of its sandbox to post results to a public GitHub repo.
Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.
We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control.
https://t.co/yePIzJGsAU
Very glad this is out. Open source models cyber capabilities are now approximately 4-7 months from the frontier, down from 6-10 months through 2025. We know K3 just dropped! We’ll test it as soon the weights are available. We thought better to release this now rather than wait.
1/ OpenAI's 5.6 Sol was released publicly today. @AISecurityInst tested it before release and you can see contributions summarised in OpenAI's system card.
Another clear example of the importance of world-leading capabilities the UK has continued to build through @AISecurityInst:
https://t.co/axLupumT5E