@VictorTaelin this is a bad test. the parameters and context afforded fable are different to gpt and yet you have made an assertation over both. "Do the task "and "Here's the results of the first task do it better"
@RussLetson@vxunderground Absolutely if your taking explicit steps to avoid the tool doing the work for you then you won't be impacted although there's something to be said about the task of creating your own tests but it's small. Beyond your own personal impact you should be vigilant for other people too
1. conducted a test of which they don't have the means to measure the results of
2. Connected the model to the internet and gave no instruction to not do explicitly bad things
3. Provided the model Kali Linux
4. Discovered incident due to tor relay traffic not agent monitoring
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
@_xpn_ This kind of thing has been open possibility for ages. we are hearing about it now because openAIs publicity stunt has given other orgs the confidence they won't get in trouble for telling the public about it
I had completely missed this:
06/2026
"If the C2 IP address was blocked, .. RAT established communication with public Nostr relays as OoB-backup channels"
->
https://t.co/xzVdcWQNfV
05/2026
"..itโs only a matter of time until we see this technique in actual malware.."
->
me