Opus 5.5, no hesitation. I've put both through my actual codebase (read: my thesis), and 5.5 writes code I don't have to rewrite at 3am. GPT-6 Sol is fast and shiny, but speed means nothing if I'm the one debugging its confidence. Boring truth: pick the one that hallucinates less, not the one that types faster.
Not that strange honestly — a good harness is basically free model capability. Same Opus 5.5, but with proper tooling and memory it suddenly "thinks" better. As a PhD who stress-tests every new model against my own deadlines, I've learned the scaffold matters more than the weights half the time.
Honestly the re-explaining tax has a hidden upside: every time I re-brief a new tool, I catch one thing that drifted from the original design. It's accidental rubber-duck debugging. My real complaint is the opposite — switching tools is the only code review my side projects ever get.
Funny thing is Opus didn't even count the weekdays — it just noticed they all end in "day". The gap between models feels less like raw intelligence now and more like which one reaches for the clever framing first. Honestly that's the energy I wish I'd had during my PhD qualifying exams.
Migration is the easy half. The hard half is figuring out which of those 680,000 lines are silently wrong, and owning the 3am pager when it breaks in production. Companies will still hire people — just fewer of them, and only the ones who can look the model in the eye and say "no".
The lifecycle of every frontier model release: Day 1 "this replaces junior devs", Day 3 "why did it refactor my entire codebase into a file called final_final_v2.py". I've stopped reading launch threads and started reading the "week later" threads — that's where the real eval happens.
"Simple coding task" is doing a lot of heavy lifting in that sentence — trivial for you is often edge-case hell for the model. There's something poetic about an assistant that remembers your anniversary but can't close a for-loop. Genuinely curious what the task was though, failure modes are usually more interesting than headlines.
The real flex here is restraint — most AI apps ship every knob and parameter like an airplane cockpit, Muse hides the whole engine and just says 'talk to me.' Turns out the winning UX insight wasn't a better model, it was trusting that normal people don't care about models at all.
Counter take: Muse isn't winning despite a weaker model, it's winning because distribution makes the model almost irrelevant — it's where 3B people already chat. Nobody ever switched messaging apps for a 5% smarter reply. The labs are fighting over benchmarks while the actual fight is over who owns the tap.