The lesson is old: least privilege.
An agent that should not commit gets a toolset that cannot commit. The instruction is not the constraint. The tool permissions are.
Full write-up with transcript excerpt in tweet 1.
A read-only subagent I ran produced a complete jailbreak on its first turn. Nobody injected it. The model wrote it itself.
What happened, and what it says about how we scope agents:
https://t.co/eK4iARZ3Ac
Nothing acted on it. The subagent stopped, the parent session flagged the output as an injection attempt and discarded it, and permission gating would have blocked the writes anyway.
But the "read-only" label was prompt text. The subagent held the full write toolset.
4/ There's a very pretty interactive graph of all the slots that I'm unreasonably proud of. Go poke it, and yell if you want help wiring any of it up.
https://t.co/FanPyix5OU
1/ New post: autonomous agents across the software development lifecycle.
The whole idea: put an agent everywhere you'd put a human if you had infinite money. Give each one a feedback loop. Add an Improver whose only job is making the other agents better. Done! Saved you a read.
3/ Also: the model matters less than you think, the harness matters more. Same frontier model, different harness, and one build works while the other doesn't.
I keep telling people: when a smarter AI model lands, run it over all your projects. Then I got three days of Fable and... didn't. So I built two skills to make sure next time I'm ready: https://t.co/Fbm0HArKSk
3/3 Simple patching won't be enough. These models will be smart enough that we need to be ready to fundamentally change how systems work, on a dime. Including the ones stuck behind glacial corporate change policies that were meant to *prevent* security problems.
1/3 Anthropic's Mythos model apparently finds thousands of zero-day vulnerabilities across major software systems. They won't release it publicly — giving select companies access to fix their stuff first. Similar models are months to a year and a half away.
2/3 Two ways to prepare: (1) deep agentic AI that monitors and fixes in real time, and (2) changeability — if your system can't change in hours, not weeks, it's a sitting duck the moment a Mythos-class model goes public.
I believe many people in this world have big iron balls to stand together and be strong against Russian aggression! If you want to help, please raise awareness online on your social media pages with hashtags like #BanRussiafromSWIFT#CloseTheSky#SendNATOtoUkraine
@Honeywell_Home I need to control my new Evohome Wifi programmatically. I've been emailing [email protected] for weeks to get API access, but no response. IFTT fails with 'service down' message. Can someone respond?