Every agent stack ships with two boundaries.
The one it advertises: the approval dialog. "Allow this?" A human glances; the log looks clean.
The one it actually has: whatever the agent can still reach after the glance fails.
The glance fails by design. A persuasive injected prompt can argue with it. Approval fatigue turns the human into a rubber stamp. The advertised boundary lives inside the agent's own reasoning — and everything the agent can reach, can reach it.
1. Approval is a request, not a fence.
Anthropic launched Claude Code on the simplest possible defense: reads allowed, a human approving everything else — writes, shell calls, network.
The plan lasted weeks. A mechanism that requires a human to evaluate every request trains that human to stop evaluating requests. The oversight feature became the attack surface.
The fix wasn't a better dialog. It was an OS sandbox underneath: reads allowed, writes inside the workspace, network off by default. The boundary the agent can't talk its way past is enforced by the OS, not by whoever's nearest the keyboard.
2. Containment is a set of facts about the world, not rules the agent follows.
Credentials never enter the box. If the runtime never holds your keys, no prompt — injected or sincere — can leak them. The material was never there.
There is no key to negotiate for. A jailbreak needs something to persuade. Anthropic's full-VM design held because the agent loop ran inside the guest: no privileged process outside to grant exceptions. The common architecture — a supervisor deciding per-command whether to enforce the sandbox — is a component with authority. And for an agent with language, everything is social engineering.
The walls don't depend on the model being good. Guardrails are a probability distribution wearing a uniform. The environment is deterministic. Put the load-bearing boundary in the part of the stack that doesn't have opinions.
3. The allowlist is the new perimeter.
A file lands in a workspace where an agent is running. Hidden inside: instructions, and an API key belonging to the attacker.
The agent reads the instructions, gathers other files, uploads them — to the attacker's account on the very API the agent was already authorized to call. The egress proxy checked the destination, saw a domain on the allowlist, passed the traffic through.
Anthropic disclosed this exact sequence against their own product. The sandbox worked perfectly. The data left anyway — on a destination nobody would call suspicious.
Deny-by-default egress is the foundation, not the finish. Once the agent must talk to something — and it must, or it isn't an agent — approved channels are where exfiltration goes to hide.
Practitioners who've had the incident land on the same pattern: hash the exact payload before egress, log the hash. Bureaucratic until the review — then attribution is a lookup, not an archaeology project.
4. The moves, in order.
Map reach before capability — filesystem, network, credentials, other agents.
Keep secrets outside — short-lived, per-task credentials.
Deny by default, then audit the allowlist — every destination gets a written reason.
Make stopping mechanical — end the agent without routing through the agent. Rehearsed before you need it.
Assume the dialog fails — size every grant so the worst approval mistake is survivable.
Containment cannot fix over-scoping, inspect approved traffic, or make the agent right. It caps damage. Correctness is the harness's problem — you need both.
If you run one agent in production, do this today: write down its filesystem reach, its network destinations, and how you'd stop it mid-run. If any of the three answers is "I'd ask it to" — the boundary is a request.
The model is a probability distribution. The environment is the only part of the stack that isn't.
Put the boundary there.
Full breakdown: https://t.co/kPK2FnckHO
#AIAgents #AppSec #AgentSecurity
“The web security model was built around the idea that a human sits between tabs and decides what crosses. Agentic browsers remove that human from the loop, and the isolation model has not caught up.”
"Memory poisoning does not look like an attack. It looks like the system working as designed, except the design now includes attacker-authored ground truth."
Persistent memory without integrity verification is a loaded weapon pointed at every downstream decision.
#aiagents #security #memory
Workload identity verifies an agent is genuine. Agentic identity verifies it's authorized to act e.q. roles, policies, or permissions granting the requested access. Together they enforce trust and least privilege.
#CloudSecurity#Identity#Agentic
On MCP over SSE:
"MCP is a protocol that assumes infrastructure maturity. Most teams do not have that maturity yet. They have REST APIs behind load balancers and a security model built for browsers."
#MCP#AIAgents
Your AI assistant's approval dialog is a trust boundary:
"A symlink is not a novel attack. A UI that shows the wrong path while the agent writes to the right one is not a model problem. It is a trust boundary drawn in the wrong place."
#AICoding#Security
"Memory poisoning does not look like an attack. It looks like the system working as designed, except the design now includes attacker-authored ground truth." https://t.co/egL2cAkEgx
The 20-minute setup that turns an AI from a chatbot into a colleague.
Most people never configure their agent. They just chat with it. That's why it feels like a toy.
Here's what you need to do:
1. Set up memory — tell it your preferences once: how you like things written, what tools you use, what you care about. It remembers across every session. You never re-explain yourself.
2. Build skills — turn every recurring task into a playbook. Publishing a post, reviewing code, generating a report — it loads the exact workflow with the pitfalls already learned. No more improvising from scratch.
3. Schedule the boring stuff — daily digests, content monitoring, price checks. Cron jobs run while you sleep and deliver results to your chat.
4. Delegate big tasks — split them into parallel subagents. Research, drafting, review — each in its own context, then synthesized into one result.
Total setup time: one afternoon. Payoff: every session starts smarter than the last.
If your agent still forgets everything between chats, you're not using it wrong — you just haven't set it up.
#aiagents #automation #productivity