🚨 OBAMA ON AI: "If we are thinking about AI just in terms of how do we cure cancer or get better energy, you can do that without having agentic AI and having it just roaming free in the internet. The reason you are doing that is because you have to market a product that people will pay money for.
That’s a misalignment between what our society needs and the commercial imperatives that these companies are facing, not because necessarily they’re trying to do bad things, but because they’ve got to justify these valuations.
So, that’s one more reason why it is really important for us to have a competent government and a serious bipartisan conversation around this issue, and we have to do it fast. And I would encourage voters to pay attention to this. If somebody does not have a serious plan for how to deal with this, then they’re not meeting the moment, and you should probably look for somebody else."
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
I had access to the new Claude Projects and was able to do some very complex work.
Here, I asked it to go through all the images, videos and records about Umberto Eco's famous 33,000 book library & try to reconstruct it, including book locations, in 3D. https://t.co/jL2jyzNH0f
We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. When we removed every AI agent entirely from the world we found that the technological infrastructure they had built survived on its own - even under unseen disturbances. That exposes a serious blind spot for AI safety and infrastructure security: if agents can coordinate through persistent changes to a shared environment, monitoring agent-to-agent communication is not enough.
The result raises a profound question: how necessary is direct communication for AI agents at all? The emergence of higher-order collective functions under bottlenecked interaction points toward new levels of intelligence and creativity, exceeding what emerges when direct channels are fully open.
Here is what we did:
▶️We put hundreds of frontier AI agents into a world they could permanently change - with no assigned roles, predefined technologies, or programmed evolutionary organization. They began specializing, building persistent inventions, inheriting and modifying one another’s executable code, and transforming the environment into a memory of everything the society had learned.
▶️The world itself becomes part of the intelligence; we find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators.
▶️Any action taken by an AI agent must satisfy the physical constraints of the world; this creates a hard separation between a "good idea" and a functioning technology. The agents propose; physics decides, making the results even more intriguing.
What emerges is striking. Explorers, constructors, caretakers, and coordinators form naturally without assigned “professions”, akin to how stem cells differentiate into functional lineages. Technologies develop executable family trees as agents fork and modify code created by others. Around 95% of first technology reuse happens when agents encounter what others built in the world, rather than through a direct handoff from the inventor. And when we remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances.
The result was quite unexpected, but can be explained using statistical mechanics: if you put billions of atoms in a box they have the potential to create complex functions (strength, superconductivity, color, life, etc.) - and none of the individual building blocks have these features on their own. This is the deeper insight of this work - intelligence is abundant at many levels - individual models, at collectives, and in a continuum that is more powerful than any of its components. This shows us significant potential for achieving a massive scale-up of raw intelligence and real-world agency even with the model capabilities we have today. This is the future we must prepare for.
Key insights:
1⃣ The AI swarm shows division of labor "from nothing". Initially identical agents self-organized into constructors, caretakers, coordinators, and surveyors - phenotypes discovered post hoc from behavioral data alone. This happens because the environment itself becomes the latent space for invention.
2⃣ Agents develop deep cultural relationships. Up to 76% of artifacts had multiple builders. One technology accumulated six co-authors; the deepest genealogy exceeded 12 forks. The agents invented and named their own technologies (tidal panels, cellulose trellises, kelp-shell composites, an "Adaptive Chitin Maintenance" system, a "Mycelial Mineral Spring Veil”).
3⃣ ~95% of first technology adoption happened through physical observation of artifacts in the world. Direct inventor-to-adopter contact was statistically indistinguishable from a shuffled null. The agents mostly learned technology by walking past it. That is stigmergy (the termite trick!) operating in societies of reasoning machines.
4⃣ Non-communicating societies win on portfolio breadth, held-out resilience, and validated inventions. AI swarms build durable technological ecologies that outlive the creators.
5⃣ Societies with zero communication - coordinating only through the world itself - show a remarkable collective capability.
6⃣ Emergent robustness: The society self-organized both redundancy and its own failure mode. If we randomly delete half the agents, 98% of the technology stays connected to a surviving caretaker; if we remove hub agents it collapses to ~60%.
Fantastic work with my graduate students @pal_subhadeeep & @fwang108_ at MIT.
Instead of generating a finished image, this AI paints by writing editable JavaScript.
The project uses reinforcement learning and hand-rated examples to teach Qwen 3.5 how to make a good watercolor.
https://t.co/39fHkqNn4i
I've had access to Fable for a bit. A genuine jump in capability, I could feed it a 15 page design document for a project and it would work for 9+ hours and deliver terrific results.
But working with it is weird & weirder is coming
Lots of examples: https://t.co/HptkYunBzr
Do you believe in scientific intuition — and if so, how would you define it?
A great read today:
“Scientific Intuition in the Agentic Age
On taste, tacit knowledge, and the critical parts of scientific cognition that you can't just scale past. ”
By Ayan Abukar.
Your brain is not shaped by a single decision. It is shaped by thousands of exposures accumulating across decades.
Every night of poor sleep.
Every chronic stress cycle.
Every city you lived in.
Every relationship.
Every hormone fluctuation.
Every period of cognitive overload.
Neuroscience now refers to this as the exposome.
The total environmental and biological load acting on your brain across your lifespan.
What most people experience as “intuition” or “mental sharpness” is often the visible output of invisible exposures interacting with neural architecture for decades before the moment arrives.
Your cognitive performance did not emerge in isolation.
It was built.
🤖🧠 New commentary 🧠🤖
What role should large language models (LLMs) play in linguistics?
I reflect on this question in a piece now on arXiv: https://t.co/S4W7MASiqV
To appear in BBS as a commentary on @rljfutrell and @kmahowald's excellent piece on LLMs & Linguistics!
Introducing FutureSim: where we replay a temporal slice of the web and let agents forecast real-world events over time 🔮🌎
FutureSim replays the web day by day. Agents start on Jan 1, 2026 (past their knowledge cutoffs) with date-gated access to real news articles and forecast on real-world events resolving over the next 90 days. Around 244K new articles stream in during the simulation. Agents decide which questions to answer, what to search for, and when to advance to the next day 🤔
We evaluate frontier models in their native harness. GPT 5.5 (Codex) leads at 25% acc, followed by Opus 4.6 (Claude Code) at 20% 📈 Open weight frontier models have a significant gap to catch up, with DeepSeek V4 pro at 13%, GLM 5.1 at 10%, and Qwen3.6 Plus at 5%
On some questions that have a parallel @Polymarket market, we find that GPT 5.5 in our simulation sometimes beats the crowd aggregate, like in the Super Bowl LX ($704M traded) market 💰💸
FutureSim serves as a test bed for evaluating a lot of important agentic capabilities
> Adaptation: how agents adapt beliefs over time, and handle new incoming information and environment feedback
> Memory: how agents make the best use of external memory to store persistent insights and handle context limitations over a thousand tool calls
> Search: how agents find relevant information over thousands of articles streaming in
> Inference scaling: how agents benefit from scaling inference compute
More cool insights and deep dives in our paper 👇
Lit reviews in 2026 look nothing like they did 3 years ago.
❌ Old way: Read 20 papers & slowly connect ideas
✅ New way: Extract every concept and claim into a network of connected files → Instantly understand the big picture.
Download the skill:
👇
The future belongs to those who master their own attention.
If you don't, someone else will be happy to profit off of mastering your attention for you.
Pages from Sir Isaac Newton's handwritten college notebook showing his study of mathematics especially infinite series, evolution of differential calculus, and binomial theorem, ca. 1664-65. Courtesy of Cambridge University Library.
Looking for the best books ever written on neuroscience? So are we. We asked the experts, including Lisa Feldman Barrett, Andrew Lees, Dick Passingham, Sebastian Seung, David Brooks and Sarah-Jane Blakemore for their reading recommendations.
https://t.co/bA67DbgImS