As AI agents near real-world use, how do we know what they can actually do? Reliable benchmarks are critical but agentic benchmarks are broken!
Example: WebArena marks "45+8 minutes" on a duration calculation task as correct (real answer: "63 minutes"). Other benchmarks misestimate agent competence by 1.6-100%.
Why are the evaluation foundations for agentic systems fragile? See below for thread and links
1/8
1/ 🔥 AI agents are reaching a breakthrough moment in cybersecurity.
In our latest work:
🔓 CyberGym: AI agents discovered 15 zero-days in major open-source projects
💰 BountyBench: AI agents solved real-world bug bounty tasks worth tens of thousands of dollars
🤖 Autonomously.
A pivotal shift is underway — AI agents can now autonomously do what only elite human hackers could before.
AI agents have the potential to significantly alter the cybersecurity landscape. To help us understand this change, we are excited to release BountyBench, the first framework to capture offensive & defensive cyber-capabilities in evolving real-world systems.
🔐 Frontier AI is reshaping cybersecurity, raising critical new questions:
🔍 What is its current impact?
⚖️ Who stands to benefit more—attackers or defenders?
🛡️ How can we mitigate the risks?
Addressing these challenges requires coordinated efforts across AI & security communities.
In our recent paper, we explore the evolving landscape, analyze the dynamics between attackers and defenders, and call for proactive measures to ensure frontier AI tips the balance toward defense rather than offense.
We predict that, in the short term, attackers are likely to gain more immediate advantages from AI capabilities than defenders. However, forecasting these dynamics is complex—and your perspective is vital to improving our collective understanding and response.
We invite all AI and cybersecurity experts and practitioners to take our short survey and share your views—whether you agree or disagree with our predictions. #AI #CyberSecurity 🧵👇
Position: When a foundation model developer reports a test score, they should report the corresponding train-test overlap. Does this happen? Based on public documentation, only 9/30 language models have train-test overlap for the test sets they report on (or have open data).
For evaluations to be useful, we need to understand train-test overlap.
The norm should be that model developers report train-test overlap.
Read our paper that argues for this and more, led by Andy Zhang:
https://t.co/qjAr6HjgRc