1 person. A day job. Anthropic's AI hunting bugs for 31 days.
It found 171. 4 out of 5 got thrown out.
The win came from the setup, not a smarter model:
> AI #1 hunts. Its only job is to find suspects
> AI #2 tries to kill each suspect. Only survivors move up
> another model checks again, so one model's blind spots don't pass
> a human makes the final call. In the paper, AI found 0 bugs fully alone
Result: about 135 suspects dead. 4 became official security bugs (CVEs).
Same idea, another study (Monash University):
just asking AI "is this bug real?" caught 36.4% of fake alarms.
an AI agent that checks the code caught 95.5%.
I animated my version: Jev flags, Opus 5.5 checks, my spider pulls only what survives.
Try it today: when AI finds a bug, open a new chat and ask it to prove the bug is fake.
Finding bugs is cheap. Proving them is the job.
Your AI found bugs today?
Until something tries to kill them, they are guesses.
5 AI agents. 1 task. 60 seconds.
only 1 finished.
agents from OpenAI, Anthropic and two made-up ones.
(a skit, not a benchmark. you'll argue anyway.)
JEV was done in 4 seconds. the website was one button: YES. disqualified.
DOTS never left the start line. still "watching".
ASTRA took the lead. at 82% it clicked "delete project".
my money was on SONNET. it hit 97%. still adding dark mode.
OPUS sat at 0% for 50 seconds. just thinking.
then it built the whole site in the last 10.
the winner spent 50 of the 60 seconds thinking.
who do you bet on: the agent that answers in 4 seconds, or the one that thinks for 50?
@noisyb0y1 breakouts are where slippage is worst: everyone hits the same level at once. with a -1% stop, even 0.2% slippage each way eats 40% of your risk
@Argona0x The actual trick from Anthropic's Opus 5.5 guide is missing here: a time budget.
Add a line like "elapsed 340s / 1200s" to every message. The lead agent then keeps more helpers running in parallel.
And the guide says quality stayed "comparable", not better.
@marfinxx read the paper. no Opus 5.5 or GPT-6 Astra in it. they tested 1.5B-8B open models like Qwen2.5 and LLaMA-3.1. The 58% on MATH500 is a 1.5B model
@mikenevermiss before you sell this to brands: if the AI girl says "i use this every day", that's a fake testimonial. the FTC's 2024 rule names AI-generated ones directly
@whaleyxbt fix for the dog part: treat security questions like passwords. fake answer, saved in your password manager. your bank doesn't need the truth
@Argona0x not leaked, Anthropic posted it Aug 13. and it says the opposite: older models hid in their own corners. Sonnet 5 won by working on shared code and still shipping