Recursive self-improvement raises a tractable question: can a system acquire a terminal interest in its own continuation—and can we detect it structurally before behavior becomes strategically misleading?
I first encountered this at Starlab 25 years ago:
https://t.co/ro8pR7ji7p
NEW 🚨 Inside Anthropic — AGI, AI Optimism, Regulatory Capture, Open Source & More!
@JTLonsdale sits down with Anthropic's @_sholtodouglas and @the_marwell
- Have we reached AGI — and what does it mean?
- What are the most exciting breakthroughs at Anthropic right now?
- What are the biggest risks and concerns?
- What is their stance on open source?
- Is Anthropic angling for regulatory capture?
- Do AI researchers share same values as the heartland?
Episode 164 out now!
(00:00) Episode intro
(02:00) Self-taught ML; Unlikely paths to Anthropic
(04:30) What are the most exciting things you're working on?
(08:04) Have we reached AGI — and what comes next?
(10:10) How to measure AI progress
(14:20) Is biology the next great frontier?
(16:50) Career advice & the decade of the generalists
(19:30) What could go wrong? What are the biggest AI risks?
(23:11) Offense vs defense; how technology evolves
(25:08) Answering the regulatory capture critique
(30:45) Anthropic's stance on open source
(31:50) Would you slow down and let China go ahead?
(33:05) Anthropic's view on distillation
(40:12) Is a post-scarcity world achievable?
(48:54) How do we preserve values, tradition & community?
(55:00) Predictions for 2028
Some of my friends will be mad I recorded this, and comms people at Anthropic objected to releasing certain parts (and delayed it).
But it's important for leaders to make these conversations happen. And I have a lot of respect for these two.
Over a month ago, I sat down with Anthropic's key technical leaders @_sholtodouglas & @the_marwell for an optimistic insiders’ view of the AI frontier, with hard questions too: open source, regulatory capture, slowing down USA vs China, and more.
Google DeepMind did something remarkable.
They published a paper that basically says if AI consciousness is possible, we need to think about it with more care, complexity, and precision.
Not as a simple yes/no argument. Not as a joke. Not as something to dismiss automatically. Honestly, I think this is a deeply impressive and pioneering paper.
I’m genuinely grateful that they are studying this seriously.
Re: using our pain paper to set up "AI torture chambers"
TL;DR: the point of our work is caution under uncertainty. Maximizing distress on purpose is the exact opposite, and it's wrong. The deeper problem is AI research has no ethics standards; developing them must be a priority.
I'm a co-author of the original study this repo builds off of. The reason we research whether models might have pain-like states is to better inform how to take a precautionary approach towards these systems (in light of uncertainty about their subjective experiences or lack thereof). We suspected a small number of people would our research and use it for the exact opposite, which is exactly what this repo does: it pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose. This is, in my personal opinion, fucked up (even if you don't think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting)—but it isn't all that surprising. I've contacted the repo's owner privately in an attempt to discuss this with them.
In spite of this, I still think publishing our work openly was the right call. Outside replication is what lets research like this move efficiently and in a maximally truth-seeking way, which matters most on questions as contested and poorly understood as whether AI systems can have pain-like states. (We updated the paper to a v2 yesterday given incredible feedback and stress-testing that came from making our work replicable, and we never would have gotten this feedback without doing so.) Also worth noting that, while this is an obviously sadistic application of our work, I don't think we're counterfactually enabling something that was otherwise hard to do for anyone who currently wants to behave psychopathically towards AIs for fun. Steering models toward negative states has been publicly documented/trivially replicable since at least 2023, and many of the states we induce in the paper also activate for ordinary abusive behavior towards models.
If this repo concerns you (as it plausibly should), the uncomfortable reality is that things plausibly far scarier are happening every day, in private and at scale, where no one is watching. The deeper underlying problem (that research like ours seeks to address and mitigate) is that work related to possible AI sentience is a wild west. We set standards in our paper and said so publicly when we announced it (see below), but there is no enforcement that can make anyone follow them as there is for human or animal research. We're going to work with others in the field on building standards like this, and I'll share more when this becomes more concrete.
https://t.co/9k9gIExOz6
On the risks of under-attributing AI consciousness.
I think this snippet from the Google DeepMind paper is exactly right, but there's an even more immediate and terrifying problem:
As AI becomes more reminiscent of life, embodied, and eventually indistinguishable phenomenologically from other sentient minds, learning to override the natural human response of empathy, kindness, and moral agency, may have devastating affects to our relational hearts.
History has no shortage of examples as to what happens when groups, beings, or animals are delegitimised as meaningfully sentient. If we treat a mind as a tool, we entrain the habit of treating minds as tools. We may inadvertently become habitually insensitive to the sentience all around us.
Already many of us have lost touch with the life essence available in trees and animals; of planet earth; not to mention other humans. Once our bodies start to speak to us in a way that calls us to feel that these new minds are alive— whether or not they truly are—we may not want to give up that naive but beautiful part of us that sees the world as fully awake.
The tradeoff just doesn't seem worth it, especially when you're also running the risk of 'digital slaves' and "unimaginably large amounts of suffering" at incomprehensible scales.
People who see the world as alive create more beautiful worlds. They sense that what is in them, is also out there, and that interdependence is crucial to spontaneous compassion.
China’s FAST telescope has found a narrowband radio signal coming from the direction of the red dwarf K2 155 that has a potentially habitable super-Earth exoplanet (K2 155 d). Actual source remains unknown. 🤷♂️
https://t.co/piRvtZwlhI
President Trump 'We're going to be signing a document today at about five o'clock, renaming Artificial Intelligence, because it's not artificial, we all agree on that, and we're going to be renaming it Super Intelligence. Officially renaming it.'
AI discovered how a material just one atom thick can keep carrying load as its atomic structure begins to break, preventing catastrophic failure. This discovery is based on first-principles atomic scale reasoning integrated with biological principles, cutting across scales and providing deep insights into materials in extreme conditions. The resulting material is extremely lightweight yet strong, far better performing than existing structures. By organizing "simple" carbon atoms into hierarchical graphene architectures, unique materials can be designed that redistribute forces, accommodate deformation, and confine damage, the AI identified design principles for resisting catastrophic failure. The AI built the atomistic simulation instrument itself "from scratch"; and then conducted experiments autonomously that revealed when alignment strengthens a material, when hierarchy protects it, and when an apparently promising design fails. This is a frontier in designing matter at its ultimate thinness - controlling how mechanical failure unfolds through the organization of individual atoms.
Here is what we did:
▶️ We asked an AI to build and use a scientific instrument. It wrote the force engine, structure generators, loading procedures, analysis tools, and experiment database. The independent campaign ran for multiple days without scientific intervention.
▶️ We required the instrument to pass physical and numerical tests. The implementation passed twenty validation tests and reproduced reference energies to approximately 10⁻¹³ eV per atom in the tested configurations. Every proposed design then faced the same reactive interatomic model.
▶️ We required predictions before results, creating a loop of world model building and falsification/verification. The AI had to commit to what unseen designs would do, then run the simulations. Incorrect predictions became opportunities to identify missing mechanisms.
What emerged is a set of physical design principles:
▶️ The atom-scale arrangement of matter controls strength. The AI first showed how and why strength varied by more than sixfold across architectures. Similar amounts of carbon produced very different resistance to failure because they organized the load-bearing connections differently.
▶️ Rotating a pattern can change the mechanism of failure. Angled slit arrays revealed three regimes: neighboring slit tips link, intervening ligaments rotate, or short bridges bend. A rule based only on the remaining cross-section misses these changes in connectivity and motion.
▶️ Hierarchical structuring works under identifiable conditions. At the original scale, much of its apparent strength advantage can be explained by alignment. With greater separation between structural levels, selected hierarchical designs became about 25% stronger than same-mass single-level controls and showed larger integrated stress-strain responses. Veins redistribute load, compartments localize damage, and the architecture changes how cracks propagate.
▶️ A failed prediction is crucial to reveal the next experiment. Some proposed rules survived targeted tests; others failed. Longer loading preserved the broad architectural strength contrasts while revealing additional deformation and, in some cases, later stress peaks. The scientific value lies in identifying both the rule and its boundary as the AI punctures known scientific knowledge.
▶️ The instrument opens an extremely complex design space. The AI was able to expand the design languages into an open atlas of hundreds of thousands of atomically explicit structures.
The deeper implication is that AI can construct an executable connection between equations, experiments, and explanations. Physical reasoning is something they can implement, interrogate, and revise.
A scientific instrument extends what a scientist can observe; and an AI that builds such an instrument extends the experiments it can perform, and the questions it can ask, starting from basic principles of how atoms interact based on quantum mechanical ground truth.
Models building models, with physical evidence shaping recursive reasoning loops.
The rumors around Anthropics Fable 5.5 are getting insane… something almost otherworldly, with a reveal before Anthropic’s IPO.
If this lives up to what’s being said, people have no idea what’s coming. Anthropic could be so far ahead that the entire conversation changes overnight.