@pigeon__s Thing is, Google is profitable, while OpenAI and Anthropic are burning money like crazy. That's probably the reason why Google aren't competing head to head right now. Delaying and improving an undercooked 3.5 Pro model was the right call, too.
Statement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National AssociationStatement on behalf of UEFA and its 55 National Associations
Opus 5 experience lately
Has anyone noticed this keeps happenin to Claude models?Opus 4.8, Fable 5, Sonnet 5 now Opus 5
Over time, the reasoning gets worse, hallucinations increase & the model gets nerfed
Why it's always Claude but not nearly as often with GPT Grok or Gemini
@kunchenguid Great post! You perfectly described the sentiment around Opus 5: impressive design, but that's it. GPT-5.6 is somewhat stuck in the middle between Opus 5 and Fable, but close enough to the latter to be a very competitive choice. And Grok isn't getting the attention it deserves.
@Targeeks@JeremyNguyenPhD Sad but true. I always get better results from basically all LLMs when I'm being friendly and encouraging, but Opus 5 needs a beating to work properly and follow instructions.
@MattLeonard1616@aliromman@thsottiaux No, but Anthropic being Anthropic, they obviously benchmaxxed it. Actually using Opus 5, it feels like it's dumber even than 4.8, so screw that benchmark.
@CahlDee@thsottiaux What's the point of a benchmark when it actively caps optimized model features, and rewards benchmaxxing for their internal harness instead? Models should be judged by how well they tackle a problem, not by how well they work under arificial constraints.
@aisearchio Simple task: improve the UI of my app. But it ended up breaking dozens of small things so badly, I just reverted to my old version. Money wasted, lesson learned.
@RileyRalmuto My favourite is "claude couldn't finish this response. try again in a moment." which usually means you've wasted tokens and usage time for nothing.
@dukeofdolma It's a flashy model, really good at design. But it has a tendency to make mistakes, break working code not even related to its actual task, and it has a shitty defensive personality, like e.g. "that's a pre-existing bug, not mine", when it had introduced that exact bug earlier.
@melvynx Opus 5 is very creative and good at designing things. But it has a nasty tendency to break stuff in my code, and then it gets all defensive and claims it's a "pre-existing bug" when I call it out.
@NicolasZu Cool, thanks for the tip! Just for fun, I tried this in ChatGPT in my browser. Told it to "create image pixel art game asset, a japanese town hall with falling leaves everywhere", followed by "use img2three" in the next prompt. Pretty neat!
@patloeber To my surprise, while I was thoroughly underwhelmed by Flash-3.6, I actually find 3.5-Flash-Lite to be an excellent upgrade. A budget-oriented low-end model, sure, but extremely fast and considerably better than 3.1-Flash-Lite, it's great for everyday non-coding usage.