We burned 11.7 billion tokens to find out which AI models can actually find vulnerabilities. 10 models, three runs each, 32 freshly disclosed CVEs to rediscover.
Open weights won this round. DeepSeek V4 Pro found the most, 28 of 32, ahead of Opus 5, Grok 4.6, and Sol. Qwen, Kimi, and GLM-5.3 right behind.
Full breakdown of all 10, by @iminurputer and @pilvar222: https://t.co/CwJO7tZnhR
👀🧵