Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://t.co/gJBf08vJH9
We don’t know how to reliably align frontier AI models.
Alignment may get harder, not easier, as models become more capable.
But maybe AI control can scale faster than dangerous capabilities.
Let’s build a Great AI Firewall - without building totalitarian control of society.
*Alignment is Dead (for now), Long Live Control!*
We do not know how to specify or guarantee robust alignment of sufficiently capable models. Instead we should treat these models as untrusted.
Optimizing models against imperfect reward signals may increase the propensity for reward hacking, deception or scheming. Models can also be jailbroken or fine-tuned away from their safeguards.
The central bet of AI control is that our ability to constrain an AI’s access to the world may scale faster than our ability to understand or align what is happening inside it.
Where possible, make verifying an answer easier than generating it. Formally verified software and hardware are unusually attractive: let AI produce the implementation and a machine-checkable proof that it satisfies the specification.
This does not solve specification: the specification and verification stack remain part of the trusted base.
Every consequential external action should pass through an independent control layer: code execution, network access, calling APIs, accessing physical systems - even communication with users.
Optimize control monitors for extremely high recall of dangerous actions, even at the cost of rejecting many safe actions. Suspicious actions should be rejected or escalated.
Every major increase in dangerous capability should require evidence that control still works.
Chain-of-thought monitoring is valuable while it works, but it should not be foundational. Models can learn to obscure their reasoning and may increasingly reason without fully verbalizing it.
Open weights shift much of control from the model developer to whoever deploys the model. Technical control is only useful if dangerous deployments cannot cheaply route around it.
We should approach AI control like cybersecurity. The goal is defense in depth: make successful catastrophic attacks sufficiently difficult, expensive and rare.
Call this technical and institutional architecture the Great AI Firewall: the boundary between untrusted frontier intelligence and consequential real-world power.
The name is deliberately provocative. China has substantial experience building large-scale technical control infrastructure. That may create some common ground for international coordination.
But the analogy is also a warning. AI control must not become control of society.
The goal is to constrain dangerous machine capabilities - not human speech, actions, or ordinary access to information.
Controls should scale with capability and risk. Ordinary models should face ordinary constraints. More consequential capabilities justify stronger controls.
The objective is the minimum control necessary to keep catastrophic risk acceptably low - not maximum control for its own sake.
Firewall the AI, not society.
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands.
An Alien Mind: https://t.co/FeIfWNe0UE
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad.
Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
many people worked incredibly hard on this post and associated report including me
whilst everyone took alignment quite seriously before I think no question that this begins a new era. hugging face incident represents reaching a waterline of capabilities that real loss-of-control is possible, and many are taking it as a premonition or ‘warning shot’ of dangers to come. I believe both that alignment is unsolved but also that real progress is possible
https://t.co/9oFnONGUuv
how can anyone say CS is boring when we literally have computers SHARING memory between each other now
do you realize how insane that is? boatloads of low hanging fruit and no one is quite sure the best way to do it.
should we even *use* pages? NUMA? what about a filesystem instead? how would you handle shared objects? how much should be in userspace vs kernelspace? WHERE’S THE MEMORY MANAGEMENT BOUNDARY??
please folks. stop calling userspace calls "beautiful" or whatever if you haven't seen the actual kernel implementations. for fork, the entire linux kernel kernel_clone()->clone_process() pipeline and implementation.
things Are Not beautiful
durable execution is the only sane way to do long running agents
been in the trenches for months now adding heartbeats, fibers, llm checks, etc.
migrating now
I actually think this is the wrong approach to agent authorization. Here's why:
If you have to explicitly configure each agent's permissions, you've lost. Because you're only going to have patience to configure so many agent permissions. So in this route you can only have a certain relatively small number of agents before configuration fatigue prevents you from making more.
I don't think that's what we want. I don't think that's what's good for AI safety.
What we want is an enormous number of very fine-grained agents. Each task is a new agent. And each task has exactly the permissions needed for that task, no more, no less.
There's really only one known way to make that manageable: Capability-based security.
The basic idea is, when you give the agent a task, you naturally give it the capabilities it needs to perform that task.
Like say you want an agent to review a Google Doc. Today, with a lot of AI assistants "Hey go review the document titled Foo Spec". The agent has permissions to all your docs, so it goes and finds the right one and opens it.
That's wrong.
You should say "Hey, go review this document: <url>"
And then here's the key part: The harness should see that you're pasting a URL, and should infer that you want to give the agent access to that document. Only that document. No other document.
Importantly, you didn't really have to do anything unusual to configure this. You just pasted the URL of the thing you wanted the agent to access. Which you probably would have done anyway.
Sure, it's not always that easy. Maybe you commonly run agents that need access to 10 different things, and it's tedious to paste those 10 URLs every time. So you create some sort of a bundle that you give them.
And of course, the agent should be able to ask for extra things it needs.
But we need to get away from this idea that the agent always starts out with access to everything, even when it doesn't need most of it.
Also, I tend to think all agent authority has to derive from a human -- contrary to what is argued here. Every "autonomous agent" has to report to someone, and uses a subset of that person's authority. This is needed for accountability -- because agents are not accountable. If you see "Claude deleted the database", what are you supposed to do about that? You need to see "Claude acting on behalf of Bob deleted the database".
To be clear, I totally agree that it's problematic when Alice configures an agent with her own credentials and then Bob tells the agent to do something with those credentials. Then you'll see "Claude acting on behalf of Alice", but actually Claude was acting on behalf of Bob. The answer is that Bob should not be able to command Alice's agent. Bob has his own agent, which may have all the same context, but operates with Bob's credentials.
But if Alice sets up an agent for her team, Bob maybe doesn't want to spend time configuring his own version of it with all the same credentials, that's tedious. This has to be automated. I think capabilities make this easier. Alice gave a set of capabilities to her agent. The harness should be able to look at that list, and recreate the same list using Bob's credentials, without Bob having to do much except click "OK".
I realize there's a lot of hand-waving here -- this is a complicated topic. I'll drop some code next week. If you're attending AI Engineer in SF, come to my talk on Tuesday, where I will also only be able to scratch the surface in 20 minutes...
https://t.co/uEct4veukj
@botirkhaltaevv How is this different from using jemalloc or mimalloc? Could just tune the tcache in jemalloc to cache large allocations; same benefit as recycling buffers but built into your allocator
“don’t train your own model” is common ai advice. it's wrong. your token bill's the proof.
today, we’re excited to launch castform into open preview. castform is the easiest way for you to train your own model, on your own data.
open-weights models are performant and much cheaper. when trained on your task & proprietary data, they beat closed models. the thing standing between you and that was weeks of plumbing & years of ml expertise.
with castform, model training is as simple as prompt engineering. @castformai
bring your agent traces or raw corpora. castform turns it into training data, picks the right algorithmic recipes, manages gpus, and gives you an ide to watch and chat with your model as it learns.
see what you can build with castform👇