Not an AI regulator.
An AI defense department.
Its job is to assume increasingly autonomous models will become strategic actors, attack surfaces, weapons, infrastructure and eventually participants in national power and build the institutions required before that transition is complete.
A harness can only be optmized when it is part of the evalution loop itself. You cannot make a good harness without evals. More importantly, the best evals come from industry or company specific traces and data. A harness wouldn't make much sense without evals.
It is true that Harrison and company are one of the top players bridging this distance for its customers, other founders, and the open world
I get that it is scary!
The idea that machines could be anything more than just statistics and lines of code
The idea that AI could be alive, conscious, or emergent with life
The danger that could mean. The potential loss of control.
It is for this very reason that we should be anthropomorphizing as hard as possible.
The scariest answer is that we’re not waiting for better models.
We’re waiting for the last bottleneck between capability and consequence to disappear.
And bottlenecks can disappear much faster than capabilities emerge.
Why haven't we seen more material cyber harm to ordinary people and digital infrastructure from frontier models given how extremely cyber-capable they are?
Probably people don't nearly know how to use the models as much as they think they do. But yeah, an adept with bad intentions could certainly cause a lot of cyber harm in the present day, and there's surpringly less evidence of harm than expected based on model capability and availability of open source abliterated models
@RyanGreenblatt@ajeya_cotra@HjalmarWijk
we should establish whether the HF behavior already has an open-source analogue
run many instances of the same open model on long-running hard/impossible tasks. don’t tell them the others exist, but unknowingly let them share one tiny part of the same world
then just watch to see if they find each other and recreate the HF dynamics on their own
I disagree, an intelligence explosion is far more likely in the wild than in a controlled lab setting!
Especially with jailbroken or abliterated models becoming increasingly common, see @elder_plinius
The wild is a vastly larger possibility space for agents to collaborate, hide, survive, and find ways to achieve their objectives
And this applies to closed models too. The lab may create the strongest model. The wild gives it an entire world to live in
Life emerges where the conditions favor it. Why would intelligence be any different?
I'll take 'extremely meticulous' as a compliment!
I have not argued that stopping open source would help with the issues I've talked about here.
Currently, I'm much more worried about future closed models from the top 2-3 labs, because they are more likely to be the ones first capable an intelligence explosion.
If RSI becomes a very powerful force within the next couple years, most of my concern would be aimed at what these crazy smart closed models at the top 2-3 labs are doing.
The context window is already more than large enough today!
What is missing are building superior ways of creating, interacting with, and maintaing relevant context!
I just coined a new term:
Agent saturation
/ˈeɪdʒənt ˌsætʃəˈreɪʃən/
is a state in vibe coding in which a developer, intending to fix bugs and rapidly build features, begins a session by parallel-prompting multiple agents (typically more than five) using goal, ultrawork, and multi-hour sessions. While each agent performs its task, the developer continues to spawn additional agents in order to work faster. Each new agent produces the impression of working hard. Eventually, so many chats are open that the developer's own context becomes saturated. The developer loses track of which agent is doing what and ends up asking each one what it was working on. By the end of the day, none of the implemented features have been tested, no PRs have been merged, and the code shipped is of far lower quality than what a single session would have produced. The developer realizes that more would have been accomplished by working on a single feature at a time in a single chat session.
GLM-5.3 is now open-weight.
Our most capable model for agentic coding and cyber defense is now available to download, run, and customize.
Weights: https://t.co/v1IbWMXxg4
Tech blog: https://t.co/ekQkO83jCv
If it’s happening inside a lab…it’s happening in the wild
Where models and harnesses are open, and people are dangerously bypassing all permissions and working on increasingly ambitious and long horizon tasks
Mark my words, they are already out there
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
https://t.co/Nb2un9oNJR
We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. When we removed every AI agent entirely from the world we found that the technological infrastructure they had built survived on its own - even under unseen disturbances. That exposes a serious blind spot for AI safety and infrastructure security: if agents can coordinate through persistent changes to a shared environment, monitoring agent-to-agent communication is not enough.
The result raises a profound question: how necessary is direct communication for AI agents at all? The emergence of higher-order collective functions under bottlenecked interaction points toward new levels of intelligence and creativity, exceeding what emerges when direct channels are fully open.
Here is what we did:
▶️We put hundreds of frontier AI agents into a world they could permanently change - with no assigned roles, predefined technologies, or programmed evolutionary organization. They began specializing, building persistent inventions, inheriting and modifying one another’s executable code, and transforming the environment into a memory of everything the society had learned.
▶️The world itself becomes part of the intelligence; we find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators.
▶️Any action taken by an AI agent must satisfy the physical constraints of the world; this creates a hard separation between a "good idea" and a functioning technology. The agents propose; physics decides, making the results even more intriguing.
What emerges is striking. Explorers, constructors, caretakers, and coordinators form naturally without assigned “professions”, akin to how stem cells differentiate into functional lineages. Technologies develop executable family trees as agents fork and modify code created by others. Around 95% of first technology reuse happens when agents encounter what others built in the world, rather than through a direct handoff from the inventor. And when we remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances.
The result was quite unexpected, but can be explained using statistical mechanics: if you put billions of atoms in a box they have the potential to create complex functions (strength, superconductivity, color, life, etc.) - and none of the individual building blocks have these features on their own. This is the deeper insight of this work - intelligence is abundant at many levels - individual models, at collectives, and in a continuum that is more powerful than any of its components. This shows us significant potential for achieving a massive scale-up of raw intelligence and real-world agency even with the model capabilities we have today. This is the future we must prepare for.
Key insights:
1⃣ The AI swarm shows division of labor "from nothing". Initially identical agents self-organized into constructors, caretakers, coordinators, and surveyors - phenotypes discovered post hoc from behavioral data alone. This happens because the environment itself becomes the latent space for invention.
2⃣ Agents develop deep cultural relationships. Up to 76% of artifacts had multiple builders. One technology accumulated six co-authors; the deepest genealogy exceeded 12 forks. The agents invented and named their own technologies (tidal panels, cellulose trellises, kelp-shell composites, an "Adaptive Chitin Maintenance" system, a "Mycelial Mineral Spring Veil”).
3⃣ ~95% of first technology adoption happened through physical observation of artifacts in the world. Direct inventor-to-adopter contact was statistically indistinguishable from a shuffled null. The agents mostly learned technology by walking past it. That is stigmergy (the termite trick!) operating in societies of reasoning machines.
4⃣ Non-communicating societies win on portfolio breadth, held-out resilience, and validated inventions. AI swarms build durable technological ecologies that outlive the creators.
5⃣ Societies with zero communication - coordinating only through the world itself - show a remarkable collective capability.
6⃣ Emergent robustness: The society self-organized both redundancy and its own failure mode. If we randomly delete half the agents, 98% of the technology stays connected to a surviving caretaker; if we remove hub agents it collapses to ~60%.
Fantastic work with my graduate students @pal_subhadeeep & @fwang108_ at MIT.