When I built computer vision models, the worst failure mode was a confident misclassification - a muffin labeled as a chihuahua. ๐ง๐
In cybersecurity, hallucinations have a very different failure mode.
We ran 162 experiments to understand what kind of context helps coding agents write secure code.
Instead, we found agents that produced convincing plans, passing tests, and hardened-looking code - while still shipping the vulnerability underneath.
Research: https://t.co/vbNZw4kAUi
Presenting this at AppSec Village @ DEF CON 34. Come say hi ๐