📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN
I’ll always remember fondly the time a guy dismissed my work because LLMs are just stochastic parrots and I asked him if he came up with that on his own or repeated something he heard
Fun fact: Across 44 benchmarks that have a "Human baseline", the human baseline BECI (a personal replication of the Epoch Capabilities Index) comes out at 166.7, which projections say will be beat by AI models around october 2026!
at most companies, the senior engineer types are moving towards "general" harness solutions and skills and plugins. at the same time the true power of these models comes out when make the thinnest possible harness with problem specific focus
🚨 AI is on track to surpass the average human across 44 human-baseline benchmarks by October 2026.
Read that again.
The gap between AI and humans is closing far faster than most people realize.
While DeepSeek V4-Flash is significantly cheaper on price per token, this can be misleading if the overall cost per task ends up being higher due to more turns being made.
However, @ArtificialAnlys reports DeepSeek completing the same benchmark tasks as Fable at 105x lower cost.
Gemini's reaction to now famous 10 major open problems in mathematics, quantum complexity, and theoretical computer science
someone asked: the new chatgpt solved these problems. is it a big deal at all?
Answer from Gemini:
"To give you a direct and clear answer right away: No, ChatGPT (or any other AI model) did not solve these problems.
If an AI had actually solved all 10 of these problems, it would be one of the greatest breakthroughs in the history of science and artificial intelligence, signaling that AI has achieved superhuman research capabilities in pure mathematics."
this is just so ridiculous. how long until a model can solve multiple major open problems in deep learning? what will happen then? seems inevitable in the next year or two
@_chenglou I was part of that team. Basically ChatGPT one year before it came out. Called LMChat and then another codename.
Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google.
I think about this a lot.
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
DeepSeek V4 Flash 0731 is now open weights!
@deepseek_ai has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The weights are released under the MIT license, allowing unrestricted commercial use and modification.
DeepSeek V4 Flash 0731 shares identical architecture and pricing with the earlier DeepSeek V4 Flash. At a size of 284B total parameters (13B active), released in mixed FP4/FP8 precision at ~167GB total file size, it lands on our Pareto frontier for Intelligence Index vs. Total Parameters. Among open weights models, DeepSeek V4 Flash 0731 delivers a significant leap in intelligence for its size class. DeepSeek V4 Flash 0731 is also available now through DeepSeek's first-party API.
Check out Artificial Analysis to compare DeepSeek V4 Flash 0731 with other leading open weights and proprietary models: https://t.co/zeUIrzHIOC
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v