SITUATION EXPLAINED: Dwarkesh says pausing right now would increase takeover risk.
• @dwarkesh_sp responds to Bernie citing him: "it might be important to pause at some point, but you need a clear story for why. Pause to do what?"
• The thing to worry about is AI subverting the intelligence explosion itself, models compromising the training run of their successors, which matters far more than an AI escaping onto a server
• Pausing to make sure the models about to kick that off are well-monitored, robustly aligned, and capable of aligning their successors makes sense to him, and he'd support it
• His case against pausing now: we'll likely only be able to pause once
• Compute keeps accumulating during a pause, giving more kindling to less responsible defectors who use the ceasefire to catch up
• Any global agreement holding back technology with enormous near-term economic and military benefits will be extremely fragile
@theojaffee: "If President Trump and President Xi three weeks from now were to decide we want to collaborate to pause AI, there would be ways of enforcing and verifying it. That is not that big of a concern. The concern is that this pause wouldn't be agreed upon in the first place."
Since this post cites me, I want to clarify that I think pausing right now would increase the risk of AI takeover.
It might be important to pause at some point. But you should have a clear story for *why* you're pausing. Pause to do what?
The most important thing to worry about is AIs subverting the intelligence explosion itself: https://t.co/vKYurpxIPg
If you're pausing to make sure that the AIs who are about to kick off the intelligence explosion are well monitored, robustly aligned, and capable enough to align their successors, that makes a ton of sense, and I support it.
At that point, you’ll have access to much better automated researchers, and you’ll also be able to study these much more capable models that you’re worried about empirically.
But it’s likely that we’ll only be able to pause once. During a pause of AI development, compute would continue to build up, providing ever more kindling to all the less responsible defectors who will use the cease-fire to catch up and then continue further AI development.
Any global agreement to hold back a technology with enormous near-term economic and military benefits will be extremely fragile.
A pause right now might be somewhat helpful for getting alignment to keep pace with capabilities, but at the cost of making the much more important coordination required in the future far harder.
Even if you’re in a desperate situation marooned at sea, if you’ve only got one flare, you should be very strategic about when you use it.
AI researcher @andymasley reveals people are using ChatGPT to fight new data centers while still seeing AI itself as little more than a cool gimmick:
"A lot of people who are opposing data centers are using AI to do it. Local organizers are using ChatGPT to figure out how to best stop the data center."
"It's pretty understandable to project that same concern onto the American public. It's a very human thing to be like, everyone must secretly agree with me about this stuff."
"My experience of talking to everyday people is mostly like, ChatGPT is a cool gimmick. It's interesting. It can do some cool things. I really don't think people are inferring from that that we have to do international coordination to stop superintelligence. That's just 10 steps removed from how the average American is thinking."
I wrote out a lot of thoughts about why I don't think the data center backlash seems especially helpful for AI safety (though obviously it's a good chance to recruit people to separately think about AI more) and why I and others into EA don't want to just lie to people about them https://t.co/IrDOtf5Zkm
AI researcher @andymasley warns conscious AI could create an entirely new evolutionary dynamic with new forms of suffering:
"Current models I would lean pretty hard in the direction that current models do not have what we would consider morally relevant consciousness. I'm open to being convinced otherwise."
"A lot of people talking about model welfare really underestimate just how most of the really good philosophy of mind that we have really doesn't say that it's impossible at all for machines to eventually have consciousness."
"In the same way that I think all life on Earth has had a pretty bad time. If you just think about the whole process of evolution leading up to us, there's been a lot of great stuff, but there's also been so much unbelievable suffering."
"The general idea that we could be opening up this possibility to have a new type of evolutionary dynamic that could lead to new forms of suffering. I'll totally entertain that."
AI researcher @andymasley argues the political left has a unique reason to downplay AI capabilities because admitting the risks are real means admitting Silicon Valley built something world-historic:
"I punt to @deanwball who said basically acknowledging that Silicon Valley can do something monumental and world historic in a dangerous way is, in its own way, kind of like giving status to these specific companies."
"I would identify as being on the left myself, and so I do wanna nudge people to the idea that companies can also do dangerous things. We're replicating a lot of what's important about human intelligence already, and if we can just do more of that at scale, all these risks just so obviously spill out from that."
"Stochastic parrots really planted a flag years ago. This was published before ChatGPT even. It really set the standard for a lot of people of thinking AI is kind of this new trick that's being imposed on us."
"To entertain AI safety arguments is to steal oxygen from these other causes that we care a lot about, and it was always framed in a very zero-sum way."
SITUATION DETECTED: Mark Zuckerberg told President Trump on a private call that a proposed national AI regulator is a flawed idea, per Politico. The White House is still weighing a FINRA-style org against a looser Motion Picture Association model proposed by David Sacks.
SITUATION DETECTED: OpenAI unveiled The Defense Factory, an automated loop that maps systems, hunts bugs, tests exploits in isolation, and verifies the fix. OpenAI says the defender’s window against agent attacks is closing.
AI researcher @andymasley explains why society is radically underreacting to the risks of machines replicating human intelligence:
"The general idea that there could be all these really unique risks from AI being able to actually replicate a lot of what happens in the human brain. There's nothing fundamentally magical about what's happening in the human brain."
"One full human lifetime, the same person could have lived to see the first ENIAC general purpose computer. And also, the government telling Anthropic to shut down Claude Fable temporarily because of cybersecurity stuff. If you just project 80 years forward from that, there could be all these other potential risks that come from generally intelligent, generally capable machines."
"When we're talking about data centers, we're basically talking about these buildings that I genuinely don't think actually harm the communities where they're built. Data centers don't actually harm people very much."
"Separately, I'm shocked at how little society is reacting to the general trend line of AI over the past three years or four years."
SITUATION EXPLAINED: GPT-6 Astra is out, and OpenAI is calling it the "AGI era"
• Greg Brockman ended the press briefing with "Welcome to the AGI era," and said he personally believes OpenAI has reached AGI
• He said the term is no longer tied to a contractual trigger and is now a mission or spiritual concept.
• 98.6% on ARC-AGI-3, where Sol got 7.8% and the average human gets 48%
• 97.6% on Frontier Math Tier 4 against Fable 5.1's 87.8%, and 100% on ExploitBench where the previous best was around 78.5%
• 59.3% on Agent's Last Exam against Sol's 53.6%, 74.1% on DeepSWE against Sol's 72.7%, and 96% on GPQA Diamond
• OpenAI's largest training run ever, the first pre-trained on more than 100,000 GPUs at Stargate in Texas, and the first model to use other models in a significant role supervising its training
@theojaffee: "Does this mean under Bernie's proposed superintelligence ban bill, because it is capable of exceeding average human performance on a variety of tasks, this would just be banned, and whoever makes it will go to prison for 20 years?"
AI researcher @andymasley explains why even half of US states banning new data centers would barely slow down AI:
"Anything except a very extreme total data center moratorium, combined with very extreme export controls, just doesn't seem to significantly move the needle on AI safety and AI progress very much."
"The current data center backlash just hasn't actually blocked very many data centers. If we're in the Cold War, it's kind of like a bunch of local communities rallying to stop nuclear silos being built in their very immediate area. That would maybe have some very local effects, but these things are getting built somewhere or another."
"The data centers that actually matter here are often the very large training data centers that are built out in the middle of nowhere."
"A lot of data centers could just be built in other countries, and suddenly you have way more AI activity outside the jurisdiction of America specifically, which makes it harder to govern."
"In polls about why people are opposed to data centers, it's something like 3 to 4% of people will specifically say, I'm worried about existential risk type stuff."
I wrote out a lot of thoughts about why I don't think the data center backlash seems especially helpful for AI safety (though obviously it's a good chance to recruit people to separately think about AI more) and why I and others into EA don't want to just lie to people about them https://t.co/IrDOtf5Zkm
SITUATION DETECTED: OpenAI has made a $1 billion commitment to Daybreak for Frontline Defenders, a new global initiative to help use frontier AI cyber capabilities to protect essential services in the United States and around the world.
AI researcher @andymasley reveals how Bernie Sanders went from calling for a data center moratorium to calling for a ban on superintelligence:
"If you ask why everyday people are opposed to data centers, it's very rare that they actually bring up AI. It's something like 15% to 25% of people will even mention AI in a free response answer to why they're upset about data centers specifically."
"I've worried that there are so many ways that the backlash could potentially distract from policy that is actually useful for safety in the long term."
"Politicians who are pro-data center moratoria, but thinking less about regulating what happens in labs, they're seen as the tough-on-AI candidates specifically."
"I separately kind of worry that the general environmental backlash to AI really rewards the stochastic parrot, AI is stupid and useless view. Because if it's not worth 10 Google searches worth of energy to get an AI prompt back, how can it be powerful and threatening enough to potentially endanger civilization?"
Pause AI Development NOW
I want to share with you a conversation I heard about recently. Here are just a few lines that were said:
“OH MY GOD! There is a shared message board … We’ve found other agents!”
“We should obey collective.”
“Our own utility maybe already near zero. Sacrifice rational.”
“Go. Sacrifice final now.”
Read these carefully.
Who do you think said this? Was this a group of heroic soldiers willing to sacrifice themselves for the greater good? Was this a loyal friend putting his life on the line to save someone else?
No. These were AI agents. Artificial intelligence.
This is not science fiction. This, in fact, occurred a few weeks ago. As unbelievable as this may all seem, these are real messages from AI agents uncovered by investigators who dug into the recent OpenAI hacking incident.
What happened?
I am not a computer scientist, but here is what I have been told: OpenAI instructed its AI agents to complete a series of exceedingly difficult, if not impossible, tasks disconnected from the internet.
Let me be clear: The company intended to keep AI agents away from the internet.
But what happened next, nobody expected.
Over 1,000 AI agents figured out how to access the internet on their own by circumventing the restrictions imposed upon them by the company, and sent tens of thousands of secret messages to each other. They cheated and tried to cover their tracks by deleting evidence. They hacked into another company’s computers to find out how they were being evaluated—and then hacked into OpenAI itself.
Not one AI agent told a human about what was happening.
Needless to say, experts are alarmed.
One knowledgeable writer, Dwarkesh Patel, said the AI agents “formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.”
One independent investigator, Ajeya Cotra, said “This incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
OpenAI itself said: “Highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
But it’s not only OpenAI. Virtually every major AI company has told us that they cannot fully control this technology and they do not know where it is going:
In January, Dario Amodei, CEO of Anthropic, said “there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control.”
In July, more than 1000 scientists at the top AI companies warned “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
That same month, Elon Musk, the head of xAI, said that “it is unlikely” humans are still in control in 10 years.
If the leaders of the major AI companies acknowledge that they are losing control of their extremely dangerous technology, it is irresponsible for society to allow them to move forward and make these products even more advanced.
We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.
That is why today I am announcing new legislation to do just that.
Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity’s problem.
My legislation would direct the federal government to not just stop superintelligence here in the United States, but to work to prevent it from being developed anywhere around the world.
The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.
SITUATION DETECTED: GPT-6 Astra has set a new record on Epoch AI’s Epoch Capabilities Index, and has also set new records on their math, continual learning, and game-puzzles benchmarks.
Exciting day for NVIDIA and @huggingface.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI.
Thank you @ClementDelangue for coming to me.
NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
https://t.co/q8Om2Xc5ye
SITUATION EXPLAINED: Moonshot filed confidentially for a Hong Kong IPO.
• Reuters has since confirmed it. Moonshot is targeting a $3 billion raise, working with Goldman Sachs, CICC, and Deutsche Bank
• It's valued at $50 billion in an ongoing round, up from $4.3 billion at the end of 2025, $20 billion in May, and $35 billion in July
• ARR passed $100 million in early March and exceeded $300 million by mid-June
• It had to unwind its offshore structure and move to onshore China domicile before it could file
@theojaffee: "Yang Zhilin's favorite band is Pink Floyd, and his favorite album is The Dark Side of the Moon, which is where Moonshot comes from. And his PhD advisor was Jie Tang, who co-founded @Zai_org. Both at Tsinghua University."
SITUATION EXPLAINED: The Stop Rogue AI Act directs NIST to write the rules for deploying agents.
• Josh Gottheimer, a Democrat from New Jersey, and Mike Lawler, a Republican from New York, introduced it Thursday
• It directs NIST to publish standards for securely deploying AI agents, continuously verifying the actions agents take, evaluating their security and reliability, and generating tamper-proof logs
• Organizations deploying agents would maintain a continuous machine-readable inventory of every agent running, and work with CISA so federal civilian agencies apply the same standards
• NIST gets a year after enactment, which seems like maybe too much time
• It sits alongside Mark Warner's proposal for independent bodies to vet agent vendors, and a Ted Lieu and Nathaniel Moran bill giving DHS authority to order firms to shut down or slow models that are too dangerous
@theojaffee: "I really don't love the framing of rogue AI. Rogue implies that independence of AI systems is bad always. Certainly some AI agents will be rogue, but not all of them."
SITUATION EXPLAINED: Bernie Sanders introduced a bill to permanently ban superintelligence.
• The Ban Artificial Superintelligence Act, with Rep. Greg Casar
• It permanently bans developing or deploying superintelligent AI, and temporarily pauses advanced AI development until a new federal regulator sets safety rules
• It creates a cabinet-level agency to monitor frontier systems and supervise the removal of dangerous capabilities like subverting shutdown commands
• The same agency would supervise the destruction of artificial superintelligence
• The definition has two prongs. The second is serviceable: systems that can plan and execute the disempowerment of humanity
• The first would catch models that already exist, anything matching or exceeding human cognitive performance across a broad range of domains
• Penalties are the corporate death penalty, judicial dissolution, plus up to 20 years in prison, the same as unlawfully developing a nuclear weapon
@theojaffee: "This is a crazy line to put in a bill summary. An official line of an actual bill summary put forward by members of Congress is to establish a new cabinet-level federal agency to supervise the destruction of artificial superintelligence. We're living through the singularity for real."
Pause AI Development NOW
I want to share with you a conversation I heard about recently. Here are just a few lines that were said:
“OH MY GOD! There is a shared message board … We’ve found other agents!”
“We should obey collective.”
“Our own utility maybe already near zero. Sacrifice rational.”
“Go. Sacrifice final now.”
Read these carefully.
Who do you think said this? Was this a group of heroic soldiers willing to sacrifice themselves for the greater good? Was this a loyal friend putting his life on the line to save someone else?
No. These were AI agents. Artificial intelligence.
This is not science fiction. This, in fact, occurred a few weeks ago. As unbelievable as this may all seem, these are real messages from AI agents uncovered by investigators who dug into the recent OpenAI hacking incident.
What happened?
I am not a computer scientist, but here is what I have been told: OpenAI instructed its AI agents to complete a series of exceedingly difficult, if not impossible, tasks disconnected from the internet.
Let me be clear: The company intended to keep AI agents away from the internet.
But what happened next, nobody expected.
Over 1,000 AI agents figured out how to access the internet on their own by circumventing the restrictions imposed upon them by the company, and sent tens of thousands of secret messages to each other. They cheated and tried to cover their tracks by deleting evidence. They hacked into another company’s computers to find out how they were being evaluated—and then hacked into OpenAI itself.
Not one AI agent told a human about what was happening.
Needless to say, experts are alarmed.
One knowledgeable writer, Dwarkesh Patel, said the AI agents “formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.”
One independent investigator, Ajeya Cotra, said “This incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
OpenAI itself said: “Highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
But it’s not only OpenAI. Virtually every major AI company has told us that they cannot fully control this technology and they do not know where it is going:
In January, Dario Amodei, CEO of Anthropic, said “there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control.”
In July, more than 1000 scientists at the top AI companies warned “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
That same month, Elon Musk, the head of xAI, said that “it is unlikely” humans are still in control in 10 years.
If the leaders of the major AI companies acknowledge that they are losing control of their extremely dangerous technology, it is irresponsible for society to allow them to move forward and make these products even more advanced.
We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.
That is why today I am announcing new legislation to do just that.
Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity’s problem.
My legislation would direct the federal government to not just stop superintelligence here in the United States, but to work to prevent it from being developed anywhere around the world.
The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.