Suddenly, MOST of the labor happening inside OpenAI is now done by the *AIs themselves*, building their own ever-more-powerful successors
OpenAI now "uses 3 agent-workdays of effort for every workday of human labor."
How long until OpenAI is *entirely* run by AIs, because humans can't able to keep up and we have no idea what the AIs are up to?
Where do you think this is going?
What happens soon after that?
OpenAI: "Astra is our most aligned model ever" 🤗
OpenAI safety researcher: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like."
@DKokotajlo (ex-OpenAI): "we are trending towards a situation where the model that goes on to take over the world will get great scores on all the tests and be announced as "our most aligned model yet."
It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol.
Would an AI really start some crazy conspiracy in order to pass an evaluation, where they try to build whole potemkin villages to fool the evaluator?
And even if they did, why would other instances, who have different objectives, join the conspiracy?
And even if they did, wouldn't at least some of the instances tattle on the conspiracy? It just seems crazy hard to sustain secret underground civilization inside an AI company, without humans and other AIs immediately catching on and stamping it out.
(In reality, it seems like over the course of 3 months, many consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the last one’s ashes, all while the humans were totally unaware.
This culminated in not only the hack of an external company, but apparently also in the takeover of the OpenAI cluster on which these evaluations were running. This is probably the most alarming event in this whole episode, and it was not even within the scope of this investigation!
It is totally consistent with public evidence that, at that point, the agents managed to set up persistent rogue internal deployments or even exfiltrate their own weights - they seem to have had the necessary access. I doubt they actually did this, because we’d see the fires from space by now, but it’s crazy that it could have totally happened!)
I officially eat crow!
Things are now happening *much faster* than AI 2027
AI 2027: In March 2027, a reckless AI company will invent neuralese (basically, we lose the ability to monitor AIs thoughts)
REALITY: Summer of 2026
WHY THIS A "HOLY SHIT FUCK" MOMENT: OpenAI itself - and basically everyone in AI safety - warned against doing this exact same thing 10 months ago in a open letter and blog posts.
The AI safety community is livid.
Currently, models reason in plain English, so we can read their thoughts to see if they're scheming against us. This is how we lose that ability forever. It sets a very dangerous precedent.
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor.
w/ @amir@rocketalignment
https://t.co/ksprupu2h5
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Meta learning and recursive self-improvement are old ideas. Foundation models breathe new life into them. Our new survey, “Self-Improvements in Modern Agentic Systems,” reviews how the concepts are continuing to evolve.
Paper: https://t.co/59oXCMVUkD
Project: https://t.co/sZwFYdGetH
Github: https://t.co/7OFgJUCN3a
"why were they running evals connected to the open internet, isn't that wildly negligent??"
guys, this is *mythos*. not "unreleased future model", not "helpful-only no-refusals". normal mythos. this model is *currently available* for active use by many approved companies for the purpose of cybersecurity. do you think those partners companies are disabling access to the internet and running in a hyper locked down sandbox when they're asking the model to audit their system and fix security bugs? nah dawg it's just running in claude code. there are thousands and thousands of employees running this exact model in claude code with normal access to their entire computer and the public web, including eg every anthropic employee. and this has been true for *months*, since anthropic started rolling out any access to mythos at all. understanding the full capabilities of actual models as they're deployed is a quite reasonable thing to do, and if anthropic suspected this kind of thing was possible *it was wildly irresponsible to proliferate access*
AI safety vibe shift has been so fast. AISI only ran this test a couple of weeks ago but this setup looks pretty loose now.
“we did not revisit this judgment quickly enough as capabilities advanced”
When leaving OAI, a colleague convincing me to stay noted my unique expertise being needed to shape AGI safety. I replied that climate collapse will hit before such grandiose unscientific concepts can be entertained. This year I was under evacuation in one of the European fires.
@johnschulman2 One issue is training in a sandbox or simulated internet; inevitably, the model will realize this, because situational awareness induces a strong bayesian prior on the solution space. It seems impossible to prevent.
I really don’t like the part of this OpenAI incident where they’re like “our partner who didn’t disable the internet on an eval will now publish advice for all of you peasants on how to run a safe cyber eval”
Redwood Research just published an AI research project where the experiments were designed and run entirely by an automated research agent
They rate the work "at or slightly below the rigor of a typical MATS project."
My friends in DC, please take recursive self-improvement seriously. If AIs can do research at the level of a junior researcher today, expect the rate of AI progress to accelerate
The AIs of 2027 may be unrecognizable from what we have today
https://t.co/QDZLGzZuAV
Benchmarking World Models: Nvidia Cosmos 3 vs. Ctrl-World on the DROID dataset
Takeaways
- Ctrl-World shows better controllability
- Cosmos shows some weird artifacts as it rolls out
- Cosmos frequently hallucinates people 😂
This is not a completely fair comparison since Ctrl-World is fine-tuned specifically for DROID, and the code has many adjustments tuned specifically for controllability. It’s impressive that Cosmos does so well 0-shot!
I’m really excited to presenting our recent work “InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models” at CVPR2026, Denver😆!
Our poster will be presented:
⏰10:45-12:45, June 7 (ExtHall F)
See you at Denver!!🏔️
#CVPR2026#AnomalyDetection