Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier
Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026
Key takeaways for the reasoning variant:
➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview
➤ A significant share of the Intelligence Index gain is driven by higher abstention on AA-Omniscience, not higher accuracy. Qwen3.7 Max's accuracy on AA-Omniscience dropped 7.6 p.p (37.7% to 30.1%), while its hallucination rate dropped 21.3 p.p (44.2% to 22.9%). The model is choosing not to answer more questions rather than recalling more facts. Because hallucination rate and accuracy both feed into the Intelligence Index, the hallucination reduction is one of the larger single contributors to the +4.8 point gain on the Intelligence Index
➤ Qwen3.7 Max used 96.7M output tokens to run the Intelligence Index, ~31% more than Qwen3.6 Max Preview (73.9M). It sits mid-pack on frontier token usage: above GPT-5.5 (high, 44.5M) and Gemini 3.1 Pro Preview (57.3M), below Claude Opus 4.7 (Adaptive Reasoning, Max Effort, 112M), Kimi K2.6 (166M) and DeepSeek V4 Pro (Reasoning, Max Effort, 187M)
Key model details:
➤ Context window: 1M tokens (up from 256K on Qwen3.6 Max Preview)
➤ Multimodality: Text input and output only
➤ Pricing: Yet to be announced (Qwen3.6 Max Preview is priced at $1.30/$7.80 per 1M input/output tokens on the @alibaba_cloud first-party API)
➤ Licensing: Proprietary, closed weights
This morning, Jensen Huang, Founder and CEO of @NVIDIA, met with Carnegie Mellon students ahead of Commencement. The visit created an opportunity for students to connect and share their work. Watch his 2026 remarks: https://t.co/zwKLse7suM
Tesla Vision allows us to deploy airbags up to 70 milliseconds earlier if your Tesla detects an unavoidable collision
This can be the difference between serious injury & walking away from a crash
New Anthropic research: Teaching Claude why.
Last year we reported that, under certain experimental conditions, Claude 4 would blackmail users.
Since then, we’ve completely eliminated this behavior. How?
Fun fact - if you have a recent commit that mentions OpenClaw in a json blob, Claude Code will either refuse your request or bill you extra money.
This is an empty repo, I'm just calling Claude Code directly. Insanity.
Qwen-Image-2.0-Pro is now live 🚀🚀
We’ve pushed image quality, multilingual text rendering, and instruction following to a new level, while making performance much more consistent across styles.🌅🌃
Ranked #9 worldwide for Text-to-Image on @arena
🔗Try it now on
ModelScope:
https://t.co/pPtrbjzzBK
https://t.co/raB6WWMEMP
API:https://t.co/EgYS5qt2bF
For @DemisHassabis, the path to AGI started in 1988 with an Amiga 500 and a game of Othello. 🕹️
His epiphany that software could act on our behalf remains at the heart of our work today as we apply the same logic to solving scientific grand challenges.
Read more on @FastCompany → https://t.co/BNm6gLWS4W
New Anthropic research: Project Deal.
We created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues’ behalf.
GPT Image 2 (high) debuts at #1 on our Text to Image Leaderboard, surpassing Nano Banana 2, FLUX.2 [max], and Seedream 4.0 in the Artificial Analysis Image Arena.
OpenAI’s latest image model is a leap forward in prompt adherence, photorealism, and text rendering. We see the biggest difference in our most complex prompts, especially where no model to date has been able to follow the instructions.
On our Image Editing leaderboard, GPT Image 2 appears to be much less of a leap forward, landing approximately in line with GPT Image 1.5. Our Image Editing leaderboard tests prompts that make changes to a single input image (eg. change text, remove an object, add a person).
GPT Image 2 (high) is priced at $211 per 1k images via API, positioning it at a higher price point than Google’s Nano Banana 2 ($67 per 1k images).
It is available via OpenAI’s developer API and in ChatGPT.
See below for comparisons between GPT Image 2 (high) and other leading models in the Artificial Analysis Image Arena 🧵
Dario is wrong.
He knows absolutely nothing about the effects of technological revolutions on the labor market.
Don't listen to him, Sam, Yoshua, Geoff, or me on this topic.
Listen to economists who have spent their career studying this, like @Ph_Aghion , @erikbryn , @DAcemogluMIT , @amcafee , @davidautor
The Gemini app is now on Mac.
With this new desktop app, you can access Gemini from any screen with Option + Space and share your window to get answers based on the documents, code, or data you're working on.
We've signed an agreement with Google and Broadcom for multiple gigawatts of next-generation TPU capacity, coming online starting in 2027, to train and serve frontier Claude models.