I tested Sonnet 5.5 and couldn't believe my eyes. 2 real repos, 105 bugs, find and fix what you can.
It took it seriously. The results:
- Sonnet 5.5 (max): 55.5
- GPT-6 Astra (max): 45
- GPT-5.6 Sol (max): 43.5
- Fable 5.1 (max): 43
- Opus 5.5 (max): 41.7
- Sonnet 5.5 (xhigh): 39
The secret? Sonnet 5.5 (max) might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns.
It took about 6× Astra's turns and about 3× Opus 5.5's for 10–14 more bugs. xhigh cut its turns by more than half and dropped to 39, below Opus 5.5 max, which used fewer turns on average:
- Sonnet 5.5 (max): 1,330 turns (n=2)
- GPT-6 Astra (max): 222 turns (n=3)
- GPT-5.6 Sol (max): 1,083 turns (n=2)
- Fable 5.1 (max): 459 turns (n=1)
- Opus 5.5 (max): 476 turns (n=3)
- Sonnet 5.5 (xhigh): 588 turns (n=1)
More effort levels for Sonnet 5.5 with n=3 drop in this thread today 🧵
@Yozla_ zatim predloži strukturu i vizuelni pravac, napravi potrebne slike preko ChatGPT Images, izgradi i testiraj sajt i na kraju ga objavi preko Sites.
@Yozla_ Želim kompletan sajt od nule; prvo me vodi kroz pitanja dok potpuno ne razumeš biznis, cilj, sadržaj, funkcionalnosti i dizajn koji želim, uključujući reference poput Pinteresta, screenshotova i drugih sajtova...