I’ve been reading the sandboxing arguments between infosec people and AI alignment folks, and I tried to summarize and referee them a bit in this post. https://t.co/pjSgW9izz9
I bought a Fable dataset from one of the top Chinese LLM routers yesterday.
With just 6TB data, I can take over 7 Chinese/CIS gov entities & 19 top Chinese firms like Xiaomi, Huawei, NIO, Minimax using SSH keys, VPN configs, Aliyun keys, GitLab tokens sent to the router.
We have to build AI that murders us, because if we don’t, China will build it first, and I don’t want to get murdered by a computer that speaks Chinese. That would be ridiculous.
I keep scratching my head at how the entire “AI for defense” conversation has been 100% about fixing code. Very myopic takes all around, as if nobody has ever had to defend real infrastructure and networks. I don’t even see the AI labs discussing this dimension of it much if at all.
I deeply care about how well AI can patch bugs in code. At the same time the craziest defense things I’ve dealt with in the real world didn’t start from 0-day initial access —instead its always things like compromised long-lived secrets and credentials. AI for defense is way bigger than patching 0-days. It’s also triaging alerts, correlating events, and recognizing incidents, autonomously. We need to figure out how to hell to defend at machine speed without breaking our own networks.
The entire "AI will be good for defence" thesis rests on answers to questions like the one addressed in the below paper, and yet the frontier labs aren't bothering to answer them.
Worth thinking about why =)
I keep scratching my head at how the entire “AI for defense” conversation has been 100% about fixing code. Very myopic takes all around, as if nobody has ever had to defend real infrastructure and networks. I don’t even see the AI labs discussing this dimension of it much if at all.
I deeply care about how well AI can patch bugs in code. At the same time the craziest defense things I’ve dealt with in the real world didn’t start from 0-day initial access —instead its always things like compromised long-lived secrets and credentials. AI for defense is way bigger than patching 0-days. It’s also triaging alerts, correlating events, and recognizing incidents, autonomously. We need to figure out how to hell to defend at machine speed without breaking our own networks.
We are in some weird galaxy where all are currently true:
- humans suck at writing secure software
- LLMs also suck at writing secure software
- LLMs however are amazing at finding bugs and writing exploits
Working with frontier LLMs for vulnerability research lately has made me feel like I woke up one day and found the solid walls of my house were all made of paper. I knew humans were not good at secure software but having fistfuls of Chrome vulnerabilities makes it very real
I keep scratching my head at how the entire “AI for defense” conversation has been 100% about fixing code. Very myopic takes all around, as if nobody has ever had to defend real infrastructure and networks. I don’t even see the AI labs discussing this dimension of it much if at all.
I deeply care about how well AI can patch bugs in code. At the same time the craziest defense things I’ve dealt with in the real world didn’t start from 0-day initial access —instead its always things like compromised long-lived secrets and credentials. AI for defense is way bigger than patching 0-days. It’s also triaging alerts, correlating events, and recognizing incidents, autonomously. We need to figure out how to hell to defend at machine speed without breaking our own networks.
ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on�� far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring