@iannuttall after 1 year on gpt models only (sometimes sprinkles of fable), i will now move on 20x claude. I love astra but limits are bit too low for me, and i think sol 6.1 is great but not opus 5.5 great, so after more than 1 year i am back with ant for now! fun times!
@bridgemindai "is GPT 6 Sol with RL" ... probably from what you read here the same could be said of opus 5.5 (opus 5+RL)...RL can make it or break it.
Take also a look here: https://t.co/LIAJZPzKlJ it is impressive
So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can.
Unlike GPT-6 Sol, which was just a nerfed GPT-5.6 Terra, this one is real. The results:
- GPT-6 Astra (max): 45 for $33
- GPT-6.1 Sol (max): 44 for $6.56
- GPT-5.6 Sol (max): 43.5 for $95.35
- Opus 5.5 (max): 41.7 for $58.53
- GPT-6 Sol (max): 29.3 for $9.33
With a model like this, the 50% cut to the $200 plan doesn't matter.
n=1. More runs and results at more effort levels dropping in this thread over the next few hours 🧵
@eikkenberg@thdxr I didn't mean to sound so negative. It is still a pretty good deal and we love what they are doing... they proved a lot of times that they listen to users, i am sure it will only be better from now on (or competition will force them :D)
@mitsuhiko I'm all for trust and confidence, totally understand you. But I believe we are a lot past the possibility of building confidence reading 10/20k lines of code per day, so I use proxies for that confidence. How the program works and feels, tons of QA, tests, ...
@NicolasZu I can relate on the laziness aspect... I also see it in astra often. I think the openAI lineup it is still great for a lot of real, difficult work. it is also clear that opus 5.5 have a something new, magical in it...man i love this competition