Opus 5 launch what should we do?
> if you using "5.6 sol max" just switch to "opus 5 high" same token, same steps. but cheaper for almost 30% discount.
Comparison:
> you ask 5.6 sol max 100 question , get free 30 question without paid. insane right?
Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2%
The previous high score (7.8%) was set by GPT-5.6 Sol (Max)
Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable
New CursorBench results are in and Opus 5 is the story.
Fable 5 Max: 70.5% at $17.32 per task.
Opus 5 Max: 70.0% at $8.23 per task.
Half a point behind. Less than HALF the price. 40% fewer tokens.
And it gets better down the card. Opus 5 Extra High beats Fable 5 Extra High outright and costs $4 less per task.
Anthropic just made their own best model obsolete on value.
This is huge.
Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents.
124B parameters. Just 5.1B active per token.
With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.
Wow i see bulish on google! congrats for the release ! @GoogleDeepMind
Almost on par performance with GLM 5.2 with 3 times faster!
so 3.6 flash is = GLM 5.2 Turbo Max.
Bravo!!!
Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors and increase token efficiency, Gemini 3.5 Flash-Lite improves by 11 Intelligence Index points while Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash
@GoogleDeepMind has released the latest updates to the Gemini model family with two new models. We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash-Lite ahead of release across Intelligence, Time per Task, and Cost per Task
Key takeaways for Gemini 3.6 Flash (high reasoning):
➤ Maintains the same Intelligence as Gemini 3.5 Flash: Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.5 Flash, and just below recently released models Muse Spark 1.1 (xhigh, 51) and GPT-5.6 Luna (max, 51). Compared to Gemini 3.5 Flash, Gemini 3.6 Flash maintains similar scores across the Index, with an improvement in GDPval-AA v2 (1421, +72) and a slight regression in HLE (38%, -3 points)
➤ Half the Time per Task: Gemini 3.6 Flash records an average time per task of 1.3 minutes, a more than 50% reduction compared to Gemini 3.5 Flash (2.7). This is driven by increased token efficiency and faster output, with speeds measured at 304 output tokens per second in our pre-launch testing
➤ Slightly lower Cost per Task: Gemini 3.6 Flash’s cost per task decreases ~18%, from $0.59 to $0.50. This is driven by lower output token use and new pricing of $1.50/$7.50 per 1M input/output tokens, down from $1.50/$9.00 for Gemini 3.5 Flash
Key takeaways for Gemini 3.5 Flash-Lite (high reasoning):
➤ Significant Intelligence improvements over Gemini 3.1 Flash Lite: Gemini 3.5 Flash-Lite scores 36 on the Artificial Analysis Intelligence Index, up 11 points from Gemini 3.1 Flash-Lite (25). This places it behind models such as Nemotron 3 Ultra (38) and DeepSeek V4 Flash (max, 40), and above Mistral Medium 3.5 (30). The biggest intelligence gains compared to Gemini 3.1 Flash-Lite are in agentic evaluations, with improvements in GDPval-AA v2 (1140, +498), TerminalBench v2.1 (53.6, +22.5 points) and Tau3-Banking (16.5%, +7.8 points)
➤ Nearly half the Time per Task: Gemini 3.5 Flash-Lite records an average time per task of 0.6 minutes, nearly a 50% reduction compared to Gemini 3.1 Flash-Lite (1.0). This is driven by increased token efficiency and fast output speed, measured at 350 output tokens per second in our pre-launch testing
➤ More expensive with 2x Cost per Task: Gemini 3.5 Flash-Lite’s average cost per task increases from $0.04 to $0.09, driven by new pricing of $0.30/$2.50 per 1M input/output tokens, up from $0.25/$1.50 for Gemini 3.1 Flash-Lite. This cost increase comes despite using fewer output tokens, falling from 20k to 13k average output tokens per task
Key model details:
➤ Context window: Both models retain the same 1M context window as their predecessors
➤ Multimodality: Both models have text, image, video, and speech input with text output only
➤ Pricing: Gemini 3.6 Flash is priced at $1.50/$7.50 per million input/output tokens, down from Gemini 3.5 Flash at $1.50/$9.00. Gemini 3.5 Flash-Lite is priced at $0.30/$2.50 per million input/output tokens, with the same input pricing across all input modalities. This is an increase from Gemini 3.1 Flash-Lite, which is priced at $0.25/$1.50 per million input/output tokens, with input audio tokens at $0.50. Both models retain the same 90% discount for cached input tokens
Open-weight models carry real responsibilities.
Hugging Face’s disclosure describes how GLM-5.2 was used in a self-hosted forensic workflow during a time-sensitive cyber incident, keeping sensitive attacker data and credentials within Hugging Face’s own environment. This demonstrates the practical value of responsibly deployed models for defensive cybersecurity and incident response.
I just spoke with the Hugging Face team last week, and we are optimistic about the path ahead. We will continue to strengthen our cybersecurity and broader safety evaluations, improve documentation and secure deployment guidance, and work with the broader ecosystem to advance the responsible development and deployment of open-weight models.
https://t.co/9emaJe12mx
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!