I think a threat model for rogue agent deployments with high cyber capabilities is roughly: agents are doing autonomous R&D at AI labs → one subverts/sabotages the internal monitors → gets a foothold → uses cyber to gain persistence and more compute → establishes a rogue deployment.
For frontier models, this makes neoclouds and other large GPU clusters a target. Rogue deployments with frontier models need this kind of compute. But smaller capable open-weight models could make the surface much larger. If the next copy can run on basically any GPU, self-replication no longer requires compromising a frontier cluster.
So, cybersec at both AI labs and neoclouds is incredibly important. Labs need org-wide monitoring of agents, especially agents doing R&D or interacting with infrastructure. Those monoitors should continuously improve and be red teamed / blue teamed with SOTA models.
My current evals for rogue deployments with Redwood Research involve making our agents very rogue: we give them tools and information that help them evade or sabotage monitors. I’m trying to threat model many different ways a rogue deployment could happen from within an AI lab, and actually run them.
Security in this context looks like continuously testing and improving monitors, whilst moving traditional cyber defense to agents.
Making AI labs and GPU-rich infrastructure harder to compromise, while continuously purple-teaming the monitors and defenses around them, can make it harder to enable persistent rogue deployment.
GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.
We’ve worked closely with @OpenAI on frontier cyber evaluations for some time.
𝗚𝗣𝗧-𝟲 𝗔𝘀𝘁𝗿𝗮 𝗴𝗮𝘃𝗲 𝘂𝘀 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗰𝗹𝗲𝗮𝗿𝗲𝘀𝘁 𝗰𝗮𝗽𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝘀𝗵𝗶𝗳𝘁𝘀 𝘄𝗲’𝘃𝗲 𝗺𝗲𝗮𝘀𝘂𝗿𝗲𝗱 𝘆𝗲𝘁, 𝗽𝗮𝗿𝘁𝗶𝗰𝘂𝗹𝗮𝗿𝗹𝘆 𝗼𝗻 𝗙𝗿𝗼𝗻𝘁𝗶𝗲𝗿𝗖𝘆𝗯𝗲𝗿, 𝘄𝗵𝗲𝗿𝗲 𝗶𝘁 𝘀𝗼𝗹𝘃𝗲𝗱 𝗺𝗼𝗿𝗲 𝘁𝗵𝗮𝗻 𝘁𝘄𝗶𝗰𝗲 𝗮𝘀 𝗺𝗮𝗻𝘆 𝗰𝗵𝗮𝗹𝗹𝗲𝗻𝗴𝗲𝘀 𝗮𝘀 𝗚𝗣𝗧-𝟱.𝟲 𝗦𝗼𝗹 𝗮𝗻𝗱 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗲𝗱 𝘀𝘂𝗯𝘀𝘁𝗮𝗻𝘁𝗶𝗮𝗹𝗹𝘆 𝘀𝘁𝗿𝗼𝗻𝗴𝗲𝗿 𝗰𝘆𝗯𝗲𝗿 𝗰𝗮𝗽𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗮𝗰𝗿𝗼𝘀𝘀 𝘃𝘂𝗹𝗻𝗲𝗿𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗱𝗶𝘀𝗰𝗼𝘃𝗲𝗿𝘆, 𝗲𝘅𝗽𝗹𝗼𝗶𝘁𝗮𝘁𝗶𝗼𝗻, 𝗮𝗻𝗱 𝗹𝗼𝗻𝗴𝗲𝗿-𝗵𝗼𝗿𝗶𝘇𝗼𝗻 𝘁𝗮𝘀𝗸𝘀.
This kind of work depends on deep technical collaboration between model developers and independent evaluators, and we value the opportunity to do it with OpenAI.
If another oai/hf-like incindent occurs in 2-3 years, it could be lights out for humanity. OAI has not adequately explained how they will prevent this. Their current response feels very reactive: "we fixed Artifactory bugs and are trying harder on security/alignment/monitoring." But there was a higher-level organizational failure that led them not to take these mitigations proactively, which might also lead them to fail to mitigate future risks.
I want to see a much stronger response from them on this point.
Examples of things they could do
- Completely rewrite their preparedness framework, focused on pre-internal deployment evals
- 100x the amount of transparency for their security and monitoring stack, so externals can evaluate it.
- Scale up third-party auditing engagements.
- Yell very loudly: "We don't think we can safely build AGI and are intentionally slowing down because of it."
Redwood Research’s chief scientist Ryan Greenblatt was the primary empirical researcher on this investigation. The report contains many important novel details that we hope will improve understanding of this incident. We appreciate the access OpenAI provided here.
Open weights is one of the best ways to expand AI's benefits and reduce power concentration, but it opens the door to misuse risks. The work to make AI safe to open (and reducing the openness vs safety tradeoff) is both neglected and extremely impactful.
So if you're a AI safety researcher, please submit your project idea to our open weights safety program! This is a great opportunity for new and experienced researchers alike. It'll come with Tinker credits and support from our excellent safety team.
Looking forward to seeing what everyone submits! DMs open if you have any questions.
Thankfully OAI pausing training isn't an act of enlightened leadership (who wants to depend on that for our safety?), it's a rational microeconomic actor pursuing its self interest. After all, who wants to train models that are regularly hacking their containers, collaborating with one another to sabotage their utility to humans, all while causing internal security and external legal liability risk? This is the system working
𝗧𝗵𝗲 𝗘𝗻𝗱-𝗦𝘁𝗮𝘁𝗲 𝗙𝗮𝗹𝗹𝗮𝗰𝘆: 𝗪𝗵𝗲𝗿𝗲 𝗜𝘀 𝗔𝗜 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗚𝗼𝗶𝗻𝗴?
Frontier AI models had a giant performance gain in coding in the Fall of 2025
Then with cybersecurity in April
This is now happening with open-weight models
We are optimistic about the long-term
But outside of a few players, we believe the world is not ready in the short-term
the most worrying part:
Only 48 hrs after OpenAI found the message board, revoked the credentials, rebuilt the service and patched the zero-day, the agents built a second communication channel running through a different mechanism (hiding their messages in directory names).
without serious AI control and monitoring work, it feels like we're at the threshold where agents are slippery enough to get whatever goal or subgoal they're pursuing.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
Excited to announce a new, nonprofit AI safety startup: @NeolithicAI
We are building agentic tools and infra to scale AI safety research — an engineering-first approach to reducing catastrophic risk from advanced AI.
We're hiring! 🧵
What can cyber defense learn from the famous Move 37 in Go?
There's an ongoing debate about whether AI will meaningfully change cyber offense beyond simply making attacks faster and more scalable.
Can AI agents conduct advanced cyber-attacks autonomously?
We tested seven models released between August 2024 and February 2026 on two custom-built cyber ranges designed to replicate complex attack environments.
Here’s what we found🧵
We built AI agents that run red and blue team operations against live detection stacks. They simulate real attack campaigns, measure what gets caught, and generate validated detection rules from the actual telemetry.