π¨ Qwen3.8 just got another upgrade.
The latest Qwen3.8 Max Preview is now live, bringing broad performance improvements and a major upgrade to the web experience.
Qwen says the model is getting better every day as development continues, with a more capable official release and an open weight version planned for the future. π
During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend.
Thank you all β the response to Qwen3.8-Max-Preview blew us away. π«Άπ«Ά
Qwen3.8 is still evolving daily. Come test it, and tell us what breaks.
We're looking forward to a more capable, official version β and to open-weight it for everyone.ππ
π¨ Claude Opus 5 is around 2Γ more expensive per task than GPT 5.6 Sol on the Artificial Analysis Intelligence Index.
Despite both models ranking among the top performers, GPT 5.6 Sol delivers similar frontier-level intelligence at roughly half the cost per completed task.
The comparison highlights the growing focus on cost-to-performance, not just benchmark scores, as frontier AI models continue to improve.
Performance wins headlines. Cost wins adoption.
Grok 4.6 is coming in 2 weeks.
Grok 4.7 follows 2 weeks later. π₯
Musk says development is accelerating as the AI model pipeline "comes together," enabling a much faster release cadence.
Grok 4.6 is expected first, followed by Grok 4.7 roughly two weeks later, with Grok 5 planned as the next major milestone.
The pace of Grok releases is about to speed up. π
Grok 4.5 is now the top-performing model for real-world invoice processing.
Ramp tested leading AI models on 150,000 invoices submitted by actual businesses, measuring whether each model could predict every correction a human reviewer would make.
Grok 4.5 achieved the highest perfect-extraction rate, outperforming similarly priced models including Gemini Flash 3.6, GPT 5.6 Terra, and Sonnet 5.
The benchmark required models to reason across 100K+ tokens of invoices, business memory, and past human corrections before processing new bills.
Zero-click accounts payable is getting closer.
π¨ BREAKING: Grok is now available inside Google Workspace for FREE.
The add-on integrates Grok directly into Docs, Sheets, and Slides, bringing AI-powered writing, data analysis, formulas, charts, presentations, and more.
No separate app. No tab switching.
AI is now built into your workflow. π
π¨ ANTHROPIC: OPUS 5 GETS BETTER AT CYBERSECURITY, BUT IT'S STILL NOT THE TOP OFFENSIVE MODEL
β’ Claude Opus 5 improves over Opus 4.8 on cybersecurity tasks
β’ Anthropic says it remains well behind Mythos 5 when it comes to developing exploits
β’ The company says Opus 5 is designed to help developers find and fix vulnerabilities, while blocking high-risk misuse
β’ The focus is on defensive security rather than maximizing offensive capabilities
This is an interesting positioning.
Rather than competing to be the best exploit generator, Anthropic appears to be drawing a line around what it wants its flagship model to optimize for.
Opus 5 is stronger than Opus 4.8 on cybersecurity tasks. But it remains substantially behind Mythos 5 at developing exploits.
Its safeguards are designed to allow developers to identify and fix software vulnerabilities, while blocking high-risk uses.
π¨ CLAUDE OPUS 5 JUST POSTED ONE OF ITS BIGGEST BENCHMARK WINS YET
β’ On ARC-AGI-3, Opus 5 scored 3Γ higher than the next-best model, according to Anthropic
β’ ARC-AGI-3 measures how well AI systems solve novel, unfamiliar problems rather than memorized patterns
β’ It's one of the benchmarks designed to test reasoning and generalization, not just knowledge recall.
If these results hold up under independent testing, this could end up being one of Opus 5's most significant performance claimsβnot because it's another benchmark, but because of what the benchmark is trying to measure.
π¨ ANTHROPIC SAYS OPUS 5 NOW LEADS ON MULTIPLE CODING BENCHMARKS
β’ Anthropic says Claude Opus 5 achieves new state-of-the-art results across several coding and knowledge-work evaluations
β’ The company is positioning Opus 5 as its strongest model yet for software engineering, research, and complex reasoning tasks
β’ It also claims the model delivers those gains while remaining highly cost-efficient
π¨ ANTHROPIC ISN'T JUST CHASING BETTER MODELS. IT'S CHASING CHEAPER ONES TOO
β’ Claude Opus 5 is designed to deliver frontier-level performance at a much lower operating cost
β’ Anthropic says it outperforms competing models while matching or beating them on cost per completed task
β’ The company is positioning Opus 5 around efficiency as much as raw intelligence
β’ Lower cost per task could make it more attractive for long-running agents and enterprise workloads.
For a long time, the question was "Which model is the smartest?"
Now it's increasingly "Which model gets the job done for the lowest total cost?"
π¨ ANTHROPIC JUST DROPPED CLAUDE OPUS 5
β’ Anthropic says Opus 5 delivers intelligence close to Fable 5 at roughly half the price
β’ The model is designed to be more thoughtful, proactive, and efficient on complex tasks
β’ Lower pricing suggests Anthropic is pushing to make frontier-level performance more practical for developers and enterprises.
The interesting part isn't just the model, it's the pricing.
Frontier performance is becoming less exclusive.
Labs are no longer competing only on capability; they're competing on how affordable that capability is to use.
If Opus 5 gets close to Fable 5 at half the cost, is that enough to change which model developers choose?
π¨ ANTHROPIC ISN'T JUST CHASING BETTER MODELS. IT'S CHASING CHEAPER ONES TOO
β’ Claude Opus 5 is designed to deliver frontier-level performance at a much lower operating cost
β’ Anthropic says it outperforms competing models while matching or beating them on cost per completed task
β’ The company is positioning Opus 5 around efficiency as much as raw intelligence
β’ Lower cost per task could make it more attractive for long-running agents and enterprise workloads.
For a long time, the question was "Which model is the smartest?"
Now it's increasingly "Which model gets the job done for the lowest total cost?"