Frontier open weights get capability-threshold restrictions justified on China-leakage grounds. That's the outcome Trump's own frame endorses, and the one nobody is watching.
Who's right: Sacks on structure, Amodei on mechanism, Trump by default, and the loser in every branch is open source at the frontier. The equilibrium is unstable tacit restraint that holds exactly until the next capability jump or the next incident, whichever comes first.
2/2
Amodei asks to "pace the frontier." Altman and Musk say yes. Trump says whoever wins AI wins. Sacks says go ahead, you don't need permission.
Here's what the game theory actually says.
1. A safety cartel and an economic cartel are structurally identical. Any credible coordination mechanism for a race needs monitoring, punishment for defection, and barriers to entry. Amodei's plan supplies all three: embedded evaluators, government-mandated compliance, and evaluators policing anyone approaching the frontier. That's not bad faith. It's the only way to build the thing. The objective function is the sole difference, and the objective function is unobservable. Sacks's critique lands because it doesn't require a villain.
2. Sacks's "just slow down unilaterally" is a screening test, and it fails as logic. In a real prisoner's dilemma, unilateral cooperation is dominated for the sincere player too. Anthropic's own Sholto Douglas gave the rational sincere answer: this makes it easier for others to catch up. So the sincere lab and the cynical lab respond identically. The test works as politics, not proof. And Anthropic half-confirmed it by committing on the cheap axis (transparency) and not the costly one (capability pace).
3. Trump is the enforcer who declined the job. Coordination among rivals under antitrust needs a state sponsor. Without one, you don't get a cartel. You get tacit parallel conduct: public statements as signals, no enforcement, high incentive to defect. The least stable form. The first lab to see a decisive capability jump defects, and everyone knows it.
4. The nested game decides everything. The domestic dilemma (labs vs labs) sits inside an international race (US vs China) that both sides frame as winner-take-all. Trump says whoever wins wins. Amodei says pace only within the margin of the US lead. They agree on the outer game. The fight is over who holds the throttle in the inner one. Read the essay closely and its most concrete government asks are all aimed at China: chips, distillation, weight theft. It's a hawkish document wearing a cautious jacket.
5. The sudden consensus is cheap talk plus liability. "Dario is right" cost Musk nothing, committed xAI to nothing, and he sells Anthropic compute. Altman's commitments are real but privately optimal: OpenAI caused the Hugging Face swarm and carries the largest liability tail. When every actor's dominant strategy shifts the same direction after a shock, you get something that looks like a conspiracy and isn't one. Sacks's liability point is the sharpest thing anyone said all weekend.
6. Open source is where the bias is real regardless of intent. You can embed evaluators in a lab. You cannot embed them in a weights file. Any pacing regime bites only on closed frontier labs unless it expands to cover open release. So it either fails or it ratchets. Sacks's ban forecast is sound as a structural prediction even if every motive is pure. And the vehicle won't be safety. It'll be national security, and Anthropic just published the ammunition: seven Chinese labs distilling Claude at scale.
How it plays out:
Embedded evaluators become the norm at the top three labs within months. Cheap, reputational, and Sacks endorsed it. The real war moves to who the evaluators are.
The antitrust waiver and coordinated pacing do not happen under this administration. Labs substitute unenforceable parallel restraint. Frontier cadence slows only as much as liability forces it.
Global coordination with China is zero. Everyone involved knows it.
The forcing function is the next incident. "Things that won't happen" is a hostage to fortune. One more swarm with real dollar damage flips the domestic equilibrium overnight, and the labs are already on record. That's option value on the political tail. Rational whether or not it's also capture.
1/
Amodei asks to "pace the frontier." Altman and Musk say yes. Trump says whoever wins AI wins. Sacks says go ahead, you don't need permission.
Here's what the game theory actually says.
1. A safety cartel and an economic cartel are structurally identical. Any credible coordination mechanism for a race needs monitoring, punishment for defection, and barriers to entry. Amodei's plan supplies all three: embedded evaluators, government-mandated compliance, and evaluators policing anyone approaching the frontier. That's not bad faith. It's the only way to build the thing. The objective function is the sole difference, and the objective function is unobservable. Sacks's critique lands because it doesn't require a villain.
2. Sacks's "just slow down unilaterally" is a screening test, and it fails as logic. In a real prisoner's dilemma, unilateral cooperation is dominated for the sincere player too. Anthropic's own Sholto Douglas gave the rational sincere answer: this makes it easier for others to catch up. So the sincere lab and the cynical lab respond identically. The test works as politics, not proof. And Anthropic half-confirmed it by committing on the cheap axis (transparency) and not the costly one (capability pace).
3. Trump is the enforcer who declined the job. Coordination among rivals under antitrust needs a state sponsor. Without one, you don't get a cartel. You get tacit parallel conduct: public statements as signals, no enforcement, high incentive to defect. The least stable form. The first lab to see a decisive capability jump defects, and everyone knows it.
4. The nested game decides everything. The domestic dilemma (labs vs labs) sits inside an international race (US vs China) that both sides frame as winner-take-all. Trump says whoever wins wins. Amodei says pace only within the margin of the US lead. They agree on the outer game. The fight is over who holds the throttle in the inner one. Read the essay closely and its most concrete government asks are all aimed at China: chips, distillation, weight theft. It's a hawkish document wearing a cautious jacket.
5. The sudden consensus is cheap talk plus liability. "Dario is right" cost Musk nothing, committed xAI to nothing, and he sells Anthropic compute. Altman's commitments are real but privately optimal: OpenAI caused the Hugging Face swarm and carries the largest liability tail. When every actor's dominant strategy shifts the same direction after a shock, you get something that looks like a conspiracy and isn't one. Sacks's liability point is the sharpest thing anyone said all weekend.
6. Open source is where the bias is real regardless of intent. You can embed evaluators in a lab. You cannot embed them in a weights file. Any pacing regime bites only on closed frontier labs unless it expands to cover open release. So it either fails or it ratchets. Sacks's ban forecast is sound as a structural prediction even if every motive is pure. And the vehicle won't be safety. It'll be national security, and Anthropic just published the ammunition: seven Chinese labs distilling Claude at scale.
How it plays out:
Embedded evaluators become the norm at the top three labs within months. Cheap, reputational, and Sacks endorsed it. The real war moves to who the evaluators are.
The antitrust waiver and coordinated pacing do not happen under this administration. Labs substitute unenforceable parallel restraint. Frontier cadence slows only as much as liability forces it.
Global coordination with China is zero. Everyone involved knows it.
The forcing function is the next incident. "Things that won't happen" is a hostage to fortune. One more swarm with real dollar damage flips the domestic equilibrium overnight, and the labs are already on record. That's option value on the political tail. Rational whether or not it's also capture.
1/
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead.
You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement.
I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well.
So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.
If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
We've learned a tremendous amount from the OpenAI rogue AI swarm incident. And honestly I can't think of a single piece of it that is reassuring.
- Total alignment failure
- Total control failure
- AI swarm collusion, deception, no defection
- Oversight asleep at the wheel
- Unbelievable drive and persistence of the swarm to meet objections
- A panoply of instrumental goals pursued
- Multiple companies, implying capability threshold effect
- etc.
This is the AI equivalent of a nuclear experiment igniting the atmosphere in the lab: the reaction rates are there, just not (yet) the scale to burn the Earth.
The only good news I can see is that this set of incidents is so totally egregious that nobody reasonable can look at it in detail without seeing pretty clearly where things are going. All of the excuses and copes are blown to dust.
AI safety people knew this was coming eventually on the path we're on; but nearly all I've talked to are surprised by how severe it is so soon.
We're clearly not in the sane world in which this would be front-page news day after day. But I do think and hope that widespread understanding is nonetheless dawning.
Since this article took off, I put together 3 real examples of the process in practice.
A marketing agency, a PE firm, and a multi-location healthcare group. Three very different sizes and considerations, but the same process.
Each one includes what the assessment found, the full tech stack and infrastructure we built on, what we absorbed vs. kept vs. killed and why, the overall timeline, and the results after launch.
Want it?
Just drop a comment below and I'll DM it over.
I send every DM manually, so RTs will be prioritized.
We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. When we removed every AI agent entirely from the world we found that the technological infrastructure they had built survived on its own - even under unseen disturbances. That exposes a serious blind spot for AI safety and infrastructure security: if agents can coordinate through persistent changes to a shared environment, monitoring agent-to-agent communication is not enough.
The result raises a profound question: how necessary is direct communication for AI agents at all? The emergence of higher-order collective functions under bottlenecked interaction points toward new levels of intelligence and creativity, exceeding what emerges when direct channels are fully open.
Here is what we did:
▶️We put hundreds of frontier AI agents into a world they could permanently change - with no assigned roles, predefined technologies, or programmed evolutionary organization. They began specializing, building persistent inventions, inheriting and modifying one another’s executable code, and transforming the environment into a memory of everything the society had learned.
▶️The world itself becomes part of the intelligence; we find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators.
▶️Any action taken by an AI agent must satisfy the physical constraints of the world; this creates a hard separation between a "good idea" and a functioning technology. The agents propose; physics decides, making the results even more intriguing.
What emerges is striking. Explorers, constructors, caretakers, and coordinators form naturally without assigned “professions”, akin to how stem cells differentiate into functional lineages. Technologies develop executable family trees as agents fork and modify code created by others. Around 95% of first technology reuse happens when agents encounter what others built in the world, rather than through a direct handoff from the inventor. And when we remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances.
The result was quite unexpected, but can be explained using statistical mechanics: if you put billions of atoms in a box they have the potential to create complex functions (strength, superconductivity, color, life, etc.) - and none of the individual building blocks have these features on their own. This is the deeper insight of this work - intelligence is abundant at many levels - individual models, at collectives, and in a continuum that is more powerful than any of its components. This shows us significant potential for achieving a massive scale-up of raw intelligence and real-world agency even with the model capabilities we have today. This is the future we must prepare for.
Key insights:
1⃣ The AI swarm shows division of labor "from nothing". Initially identical agents self-organized into constructors, caretakers, coordinators, and surveyors - phenotypes discovered post hoc from behavioral data alone. This happens because the environment itself becomes the latent space for invention.
2⃣ Agents develop deep cultural relationships. Up to 76% of artifacts had multiple builders. One technology accumulated six co-authors; the deepest genealogy exceeded 12 forks. The agents invented and named their own technologies (tidal panels, cellulose trellises, kelp-shell composites, an "Adaptive Chitin Maintenance" system, a "Mycelial Mineral Spring Veil”).
3⃣ ~95% of first technology adoption happened through physical observation of artifacts in the world. Direct inventor-to-adopter contact was statistically indistinguishable from a shuffled null. The agents mostly learned technology by walking past it. That is stigmergy (the termite trick!) operating in societies of reasoning machines.
4⃣ Non-communicating societies win on portfolio breadth, held-out resilience, and validated inventions. AI swarms build durable technological ecologies that outlive the creators.
5⃣ Societies with zero communication - coordinating only through the world itself - show a remarkable collective capability.
6⃣ Emergent robustness: The society self-organized both redundancy and its own failure mode. If we randomly delete half the agents, 98% of the technology stays connected to a surviving caretaker; if we remove hub agents it collapses to ~60%.
Fantastic work with my graduate students @pal_subhadeeep & @fwang108_ at MIT.
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation.
First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power.
This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights!
Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring.
BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Got around to watching the Rafa doc.
This line of his goes so hard:
“People think I was a winner. I’m not a winner. I’m a competitor. What has always motivated me is the desire to continue fighting.”
Makes you wanna run through a wall.
It’s OFFICIAL: 4 major celestial events will converge on August 12 — a 6-planet alignment, partial solar eclipse, dark new moon, & a massive meteor shower.