Dan Selsam’s statement is one of the more compelling arguments for AI risk I’ve read.
But I think it also reveals a huge political blind spot in the AI safety conversation.
The argument is repeatedly framed around humans losing control.
The obvious question is:
Which humans?
Most people do not control whether wars happen, whether borders close, whether they can afford healthcare, whether they lose their job, whether their currency collapses, or what laws they live under.
Some live under governments they cannot meaningfully influence or even safely leave.
If you are poor, sick, imprisoned, stuck under an authoritarian government, or simply born in the wrong place, the existing human-controlled world may already exercise enormous power over you.
So “we must preserve human control” can sound very different inside an AI lab than outside one.
To a researcher, the status quo contains enormous option value that might be irreversibly destroyed by an uncontrollable AI.
To someone getting crushed by the status quo, AI may instead represent an option to change it.
Better medicine.
Cheaper expertise.
More individual leverage.
More wealth.
Perhaps even better governance.
That person can understand the existential-risk argument perfectly well and still think:
“Why exactly should I preserve this system?”
This is not an argument for releasing an unaligned ASI.
It is an argument that AI safety has a political problem.
You cannot ask billions of people to accept slower progress by telling them that “humans must remain in control” while ignoring how radically unequal human control already is.
And this matters politically.
A politician focused on national power, economic growth, or competition with China hears “slow down AI so humans remain in control” very differently from an alignment researcher.
So does an ordinary person whose life might be radically improved by better medicine, cheaper intelligence, or greater economic leverage.
If the safety case sounds like a small group of extremely powerful people asking everyone else to preserve a system from which that group benefits enormously, it will lose.
The objective should not be:
preserve human control.
It should be:
increase the agency of ordinary humans while preventing any actor — government, corporation, individual, or AI — from acquiring irreversible control over everyone else.
Exit.
Contestability.
Reversibility.
Pluralism.
The ability to say no.
That is a much stronger thing to preserve than the fact that the entity at the top happens to be made of meat.
And there is a darker implication here.
If AI safety fails to offer people a future substantially better than the present, some people may eventually support taking the risk knowingly.
Not because they failed to understand the alignment argument.
Because they understood it and decided they had less to lose than the people asking them to stop.
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: https://t.co/TxMNr0vhrL
This is exactly what I mean by “show me the project.”
The base model “didn’t work very well”, so they built the missing value around it: domain data, fine-tuning, workflow, evaluation.
Model capability matters, obviously. But the leaderboard matters a lot less once you start building actual systems around the model
Even years ago, when the models were dramatically weaker, we were already getting absurd value from them. We were trying to generate code with early Copilot while learning how transformers worked.
They were dumb. They were useful.
And honestly, what I want most now is often not another benchmark win.
More compute. More limits. More experiments.
Low reverence, high utilization
One thing I wish I saw more in AI discourse:
A model can still be dumb and extremely useful.
I know people who will happily say “this model is still kinda dumb” - and then use it all day for real work.
“This is still dumb” and “I want 10x more of it” can both be true
One of my friends built a harness around models a long time ago. At this point he barely cares which model is winning Twitter this week.
Models are components in a system.
Planning, execution, critique, verification, retries, routing - use different models where they make sense.
Once you work this way, “worse than Astra at this demo” stops meaning “useless.”
Funny thing: I’m annoyed at OpenAI today, talked to a Dot for a bit, and it was so relentlessly positive that I started feeling a little bad for it.
My empathy apparently does not respect ontology
My attempt to steelman Tibo before DevDay:
I think OpenAI is rebasing the plans around a much higher 1X floor and moving a meaningful share of the value into features/agents that don’t consume the old usage budget in the same way.
So Pro 200 may genuinely let many users get more work done while still giving them less raw frontier-model compute than before.
If that’s what ships, Tibo’s messaging will make a lot more sense.
If not, “everyone gets more” is going to be a very hard claim to defend.
@thsottiaux If the absolute compute is going up, just publish the absolute limits. A 10X multiplier against a moving 1X baseline doesn’t tell users much.
The only question that matters is: does new Pro 200 provide more actual usable compute than old Pro 200?
A recurring failure mode in AI product communication:
“We’re giving you less, but actually you should feel like you’re getting more.”
Just say you can’t sustainably provide the old limits right now. That’s completely understandable.
You don’t get to decide what counts as “more value” for the user.
Hi,
Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan.
Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago.
(a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want.
(b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions.
(c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent.
(d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet.
I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news.
Codexingly,
Tibo
Martin Fowler wrote that he finds interacting with LLMs somewhat uncanny https://t.co/LNeasaX04H. It made me realize I experience almost the exact opposite.
People often describe talking to LLMs as uncanny.
For me it’s almost the opposite.
Humans have never had neutral access to knowledge. We learn through parents, teachers, professors, bosses, colleagues, reviewers — people with limited knowledge, ego, status incentives, moods, and sometimes a real incentive to keep knowledge scarce.
I’ve had people withhold knowledge because being the only person who knew something made them valuable.
I’ve had a reviewer leave comments essentially saying “only an idiot could have written this.”
I’ve dealt with people who made asking questions feel expensive.
Against that baseline, an LLM feels remarkably normal.
It doesn’t become less valuable when I understand what it knows. It doesn’t need me to remain confused. It doesn’t get irritated because I ask the same question five times. It can explain something differently, adapt to how I think, and after a long enough conversation start speaking in a language that feels unusually natural to me.
Of course it isn’t neutral. It inherits biases from training data, post-training, product decisions and the people building it.
And the risks are real: cognitive offloading, dependence on a persuasive mediator, losing the ability to notice when it is confidently wrong.
But the baseline isn’t an independent human reasoning directly from reality.
The baseline is a human who has always depended on other humans and institutions to learn.
Maybe Fowler and I simply have different priors here. My experience with human knowledge intermediaries makes it difficult for me to experience the machine as uncanny.
If anything, I find it surprisingly canny.
I’m asking what we tell people who are still studying, still adapting, still paying rent, and still being required to prove their competence while the meaning of competence keeps shifting.
“Don’t defend the activity. Defend the value” is good advice.
But AI may eventually produce the value too.
Then the problem stops being “how do I stay employable as a software engineer?” and becomes something much larger:
How do humans retain economic security, agency, motivation to learn, and some idea of what they want to do with their lives in a world where intelligence itself is abundant?
That seems like a more useful conversation than telling the meat to hit Enter harder.
I actually agree with a lot of this essay.
Software engineering has no sacred right to remain a human profession. If AI can do implementation better, fine. If it eventually does architecture and judgment better too, moving up the org chart won’t save us either.
But “meat proxy” turns several very different problems into one convenient insult.
I’m increasingly allergic to claims like “Fable is smarter than you, me, all of us.”
Smarter at what? Over what task distribution, on what horizon, with how much context, supervision and verification? Compared to which engineers?
I’m happy to believe a model outperforms me at plenty of things. But “it feels insanely smart” is not an eval, and deleting apparently dead code from an unfamiliar repository is not a comparison with the person responsible for keeping that system running.
A capable engineer struggling to adapt isn’t necessarily stupid, lazy, or unwilling to think. And someone running six agents at once isn’t necessarily doing good engineering.
I’ve seen AI used carelessly enough that I wouldn’t want to employ someone working that way. My objection isn’t that they delegated the work. It’s that delegation doesn’t remove responsibility for the outcome.
Adaptation speed, engineering competence, and responsible delegation are different things. “Meat proxy” doesn’t tell us much about any of them.
I know a genuinely bright engineer who is struggling to find work and whose move into AI hasn’t been straightforward either. That doesn’t prove anything about the whole labor market. It does make “just adapt” feel remarkably incomplete.
Unfortunately, the carbon-based component still needs to eat and pay rent.
And this transition can pull in two directions at once: AI increases what one person can accomplish while also raising what employers expect that person to know, supervise and take responsibility for.
Cheaper implementation doesn’t necessarily mean an easier path to employment.
You may still need systems knowledge, domain expertise, security awareness, testing, architecture — and now the ability to build reliable workflows around agents.
Imagine an AI-written job description for a role producing AI-generated software, with the supposedly obsolete human still expected to validate the whole thing.
That isn’t inherently contradictory. Delegation and verification require different skills.
But if human verification remains part of the workflow, developing that competence has to be part of the story.
You don’t explain that transition by calling its participants meat.
And employment isn’t even the part I find most confusing.
I’m experiencing another version of this as a student.
I’m finishing a Master’s final project in mechanistic interpretability. The assessed writing has to be my own, not something I hand off to a model.
Meanwhile, the models keep improving, and I keep reading that they’re smarter than all of us and that personally doing work a model could do is no longer the point.
Sometimes it feels like lifting the same heavy weight while everyone else keeps relabeling it as lighter.
It’s still heavy for me. What seems to be shrinking is the value people assign to the effort.
That doesn’t make learning worthless. I still want to understand the systems I’m studying. But it makes the question of what to learn, and why, much harder.
The problem isn’t simply that my tutor knows more than me. Tutors usually do.
It’s that the relationship between learning, producing work and getting hired is shifting while I’m still completing the qualification.
University asks me to demonstrate independent understanding. Work increasingly asks me to demonstrate effective delegation. The public conversation tells me that even the judgment involved in both may eventually be automated.
Those expectations aren’t necessarily incompatible. But connecting them takes more thought than “don’t let your brain rot.”
What should I internalize? What should I delegate? How do I develop enough understanding to evaluate AI assistance without spending all my time reproducing work the tools can already do?
And what exactly should students take away from “the model is smarter than all of us”? A learning strategy, or just a new authority to defer to?
There’s also a practical asymmetry in advice like “hit Enter and go play with your kids.”
A job isn’t merely an interesting activity plus a paycheck. For many people it is also healthcare, housing, credit, residency, immigration status, geographic freedom, and the ability to survive a bad year.
I’m living outside my home country without permanent residency. For someone whose ability to stay somewhere can depend on employment, displacement isn’t simply a career change. It can mean moving your entire life.
Being willing to let a profession disappear doesn’t mean being financially or legally equipped to survive its disappearance.
That isn’t an argument against automation. I’m not particularly interested in preserving programming jobs for the sake of preserving programming jobs. If software engineering disappears, fine. There are plenty of other interesting things to do.
It’s an argument against treating the transition mainly as a psychological problem suffered by people emotionally attached to typing code.
To the author’s credit, the essay eventually acknowledges that the demonstration produced a draft PR and Jira epics without establishing that anything important had improved.
That’s one of the most interesting parts of it.
Producing artifacts, improving a system, and helping someone become effective are not the same achievement.
And Uncle Bob: this is where I’d genuinely like to hear your own take.
After decades of teaching engineers why testing, design, discipline, feedback and responsibility matter, what does the next chapter look like?
Changing your mind about how software gets built isn’t the problem. I’d be more concerned if your advice never changed.
But which principles should students still practice themselves? Which should become automated checks? How do we teach responsible delegation? How do we assess understanding without confusing it with either typing speed or the ability to forward a convincing model response?
A repost obviously doesn’t tell me your whole position.
But there is a small irony in the “meat proxy” conversation making me want to ask for less forwarding and more of your own judgment.
I don’t expect anyone to produce an AI-proof career plan. And I don’t need reassurance that humans will always be better at something economically valuable. Maybe we won’t.
Since you’re doing an AMA, I’m still very curious about this one:
You joined Anthropic in May after they’d been trying to recruit you for ~2 years, then resigned only a few months later.
What specifically changed during that time? New evidence, capabilities, internal experience - or did your interpretation of the same facts change?
@__nmca__ “Take him extremely seriously” is not an argument either.
I actually think Selsam makes one of the stronger AI-risk cases I’ve read. But I’d rather discuss the mechanism and evidence than keep stacking impressive credentials on top of the claim.
The word “we” is doing a lot of work here.
Humanity doesn’t collectively control the future today either. Power is already concentrated among particular governments, companies, labs and individuals.
I think your second risk may actually be the deeper one: not simply human vs AI control, but whether power remains contestable, reversible, plural, and possible to exit.
@DKokotajlo This is one of the stronger AI risk arguments I’ve read, but I think it exposes a political blind spot in the safety conversation:
“preserve human control” immediately raises the question — which humans actually have control today?
I wrote out why I think this matters:
Dan Selsam’s statement is one of the more compelling arguments for AI risk I’ve read.
But I think it also reveals a huge political blind spot in the AI safety conversation.
The argument is repeatedly framed around humans losing control.
The obvious question is:
Which humans?
Most people do not control whether wars happen, whether borders close, whether they can afford healthcare, whether they lose their job, whether their currency collapses, or what laws they live under.
Some live under governments they cannot meaningfully influence or even safely leave.
If you are poor, sick, imprisoned, stuck under an authoritarian government, or simply born in the wrong place, the existing human-controlled world may already exercise enormous power over you.
So “we must preserve human control” can sound very different inside an AI lab than outside one.
To a researcher, the status quo contains enormous option value that might be irreversibly destroyed by an uncontrollable AI.
To someone getting crushed by the status quo, AI may instead represent an option to change it.
Better medicine.
Cheaper expertise.
More individual leverage.
More wealth.
Perhaps even better governance.
That person can understand the existential-risk argument perfectly well and still think:
“Why exactly should I preserve this system?”
This is not an argument for releasing an unaligned ASI.
It is an argument that AI safety has a political problem.
You cannot ask billions of people to accept slower progress by telling them that “humans must remain in control” while ignoring how radically unequal human control already is.
And this matters politically.
A politician focused on national power, economic growth, or competition with China hears “slow down AI so humans remain in control” very differently from an alignment researcher.
So does an ordinary person whose life might be radically improved by better medicine, cheaper intelligence, or greater economic leverage.
If the safety case sounds like a small group of extremely powerful people asking everyone else to preserve a system from which that group benefits enormously, it will lose.
The objective should not be:
preserve human control.
It should be:
increase the agency of ordinary humans while preventing any actor — government, corporation, individual, or AI — from acquiring irreversible control over everyone else.
Exit.
Contestability.
Reversibility.
Pluralism.
The ability to say no.
That is a much stronger thing to preserve than the fact that the entity at the top happens to be made of meat.
And there is a darker implication here.
If AI safety fails to offer people a future substantially better than the present, some people may eventually support taking the risk knowingly.
Not because they failed to understand the alignment argument.
Because they understood it and decided they had less to lose than the people asking them to stop.
@emroynoire@EthanJPerez Then I think we mostly agree. Leaving may make sense personally, but that’s different from it being an effective intervention. My question is about the latter.
Converting logic into code and memorizing tools was never the valuable part of software engineering. AI just made that painfully obvious.
The valuable part is framing the problem, choosing abstractions, reasoning about tradeoffs, and verifying the result.
And if AI gets better at that too, then sure - software engineering as a profession disappears.
The profession was never the point.