I'll be mentoring for @MATSprogram this winter alongside my colleague Sambhav. We will focus on modelling the threats from misaligned & compromised AI agents in national security deployments and on making control sketches to prevent them.
General application due on Sept 6.
Last month, the FCC added import controls on new foreign-made advanced robotics. My IAPS fellow Rocio explains clearly what this rule does, why it matters, and what still needs fixing. I'm usually an import-control sceptic, but the post convinced me that this is prudent policy!
On July 28, the FCC restricted imports of new foreign advanced robotics over unacceptable supply chain and cybersecurity risks. More is needed: component provenance checks, military cybersecurity testing, and robot model evals.
Read my explainer on IAPS' Attack Surface below.
@DaveRBanerjee Could well be a case of: https://t.co/qF7ZqMsE5S . You can get discontinuous jumps on benchmarks if there is a discontinous metric (eg matchin 99% of tests gives a score of 0, while 100% gives a score of 1)
Costs for Opus 4.5 & 4.6, DeepSeek and GLM are taken from the AISI post. The rest is estimated. Absolute numbers are sensitive to the amount of tokens that are inputs vs outputs and how much prompt caching is being done. Relative comparisons are more reliable.
I had Fable plot the per-$ performance on AISI's "The Last Ones" cyberrange. The cost-performance Pareto front consists of DeepSeek-V4-Pro + GPT-5.6-Sol.
On our cyber range "The Last Ones", GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5, from ~7 months before it.
I had Fable plot the per-$ performance on AISI's "The Last Ones" cyberrange. The cost-performance Pareto front consists of DeepSeek-V4-Pro + GPT-5.6-Sol.
I'll be co-mentoring a project in the Heron AI fellowship - a 3-month project for cybersecurity professionals to transition to AI security & safety. In our project, we'll build and evaluate honeypots that catch and incriminate malicious internal agents.
🚨New Attack Surface post: To counter AI-enabled cyberattacks, defenders and policymakers need to understand AI’s offensive capabilities and how attackers are using it. But this threat intelligence is fragile and will likely deteriorate as AI improves. 🧵
Governments should support the collection and sharing of unbiased threat intel on AI cyberattacks. They should fund AI-specific threat intel methods and cyber evals that scale with model capabilities. An Agentic Cybersecurity Exchange could facilitate unbiased intel sharing.
🎉Excited to share our #ICML paper!
🦀Open-ended and self-evolving AI systems are moving fast. Agentic systems like OpenClaw show how quickly this space is moving, with rapid capability growth 📈 and major investment💹. But safety research hasn't kept pace. 1/n
@robertwiblin I'm also confused that the numbers for GPT 5.4 don't match up with the ones they reported in the AISI Mythos evaluation. Makes me wonder whether some parts of the experiment are different https://t.co/ESM9xYCD3c
@NateBurnikell@AISecurityInst Are the expert-level narrow cyber tasks the same as CTF challenges you ran on previous models? Eg here: https://t.co/sb8EZH1Dcn
I assumed yes, but the numbers for GPT5.4 don't line up