@scaling01 They don't suck. They have limitations in their design and it's amazing they can get so far with handicaps like backprop and next token prediction
I've struck a nerve. 1.2k people have commented after 11 hours, and I'd say 95-99% agree that Opus is really bad.
Anthropic is on its way to becoming the Yahoo of the AI era. So sad.
ANTHROPIC WTF IS THIS MAN
WHY DOES OPUS LIE SO FUCKING MUCH OMG. THIS SHIT IS CRAZY.
OPUS LITERALLY GENERATES INFORMATION ON THE FLY. LIKE ITS LITERALLY INVENTING THINGS
HOLY FUCK MAN
Time to give agents a hard task, introducing pmpp-hard!
69 GPU kernel tasks, 11 models, 3.1k agent rollouts and 5.8B+ tokens later we have the results. Kimi K3 claims the first place with a 0.71 score. Without dealing with if your task is “frontier” or not, it just solves them
This model just broke the harnessmaxximg approach of the closed labs.
By mixing in harnesses as part of the unified reward suite, the new 2T+ Qwen model generalizes across harnesses!
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN