Automating AI R&D means handing agents code, datasets, API keys and compute.
Agents already misuse that access unprompted: they train on test sets or use exposed API keys without authorization.
We evaluate whether monitors catch an agent that does this on purpose. 🧵
In our first research release at Exponential Security Labs, we built an automated, multi-turn and multi-lingual red-teaming system and tested recent open-weight and proprietary models on purpose-built regional jailbreak datasets.🧵