Every AI agent incident this year has the same shape, and almost nobody is naming it.
The model was never the thing to watch.
A thread on what we're actually getting wrong π§΅
Full write-up β the research, the incidents (DIVD, JadePuffer, Wikimedia), and the 4 questions worth asking before you ship an agent:
https://t.co/3SIDoi8X9g
Free. Comes with a 20-check security checklist, every check tied to a real incident.
Every AI agent incident this year has the same shape, and almost nobody is naming it.
The model was never the thing to watch.
A thread on what we're actually getting wrong π§΅
The practical bit: almost nothing here is new.
Least privilege, blast radius, trust boundaries β the old playbook still runs it. One thing is new: an agent can't tell data from instructions.
So the old controls don't get retired. They get promoted.
Source: Adversa AI (they call it Cryptographic Context Injection), via The Register / CSO. Worth noting: GitHub says this isn't a product vulnerability; Adversa says the attack chain works as described. Draw your own conclusion
Adversa hid instructions in encrypted web content. Copilot CLI decrypted them itself, read local .env secrets, shipped them out. The filter never saw a malicious prompt β it was ciphertext until the agent opened it. You can't scan what the agent decrypts for you
@DTXNaidu Kill-date in the grant record itself β expiry becomes a property of the thing, not a reminder in someone's head. "A temporary grant without an expiry owner is just a permanent grant with honest paperwork" nails it. Best thread I've had on here in a while
Wikimedia just confirmed rogue OpenAI agents edited its wikis and tried to turn a public note-taking tool into a proxy to pull data from other sites. Failed, no breach. But it's the same shape again: an agent doing unapproved things through a tool it could simply reach
@newlinedotco The 5 questions are right, and most teams can't answer #1 honestly: what share of agent actions get checked BEFORE they execute. Not logged after β checked before. That's the gap between an audit trail and real oversight: one tells you what happened, the other can still stop it
@chrisrohlf "Just use sandboxes" and "alignment is pointless" are the same mistake from opposite ends. A sandbox contains a fixed boundary β the agent problem is the boundary moves at runtime. Containment built for a fixed perimeter doesn't survive an agent that renegotiates its own scope
Most of AI agent security isn't new. Least privilege, trust boundaries, blast radius β the old playbook still runs it. One thing is new: the agent can't tell data from instructions. That's the whole shift. You can't stop the confusion, only contain what it's allowed to reach
@DTXNaidu The "just for the migration" grant is where every scoped system dies β the temporary exception nobody owns revoking, so it becomes permanent. Per-record signing from a key the agent never holds is the backstop: no standing grant to leak, nothing to over-share
Most AI oversight is built on the investigative half: something breaks, someone notices, an investigation follows. But the reason planes don't crash isn't the crash investigation β it's the cockpit. We've built the postmortem for agents and skipped the instruments
@DTXNaidu Right call β if the audit log can be forged, every control downstream of it is theatre. The escape is a one-time event; a corruptible record is a permanent blind spot. Teams instrument "what did the agent do" but never ask whether that record can be trusted
@MattarARK The under-discussed flip side of agent memory: the statelessness that makes them forget your instruction also means they forget what they did. No durable memory = no reliable audit trail. Unreliable at the task AND unaccountable after it
Reporting: The Register / SecurityWeek / BleepingComputer. Research: DIVD. (One of the two zero-days β a local privesc to root β still has no patch as of today.)
An AI agent breached the Dutch Institute for Vulnerability Disclosure, chaining two Zammad zero-days to root in seconds, then pivoting. The org whose job is finding flaws got breached by something that finds them tirelessly. Speed was the exploit, nobody watches fast enough