DeepSeek (@deepseek_ai) opened an internal beta of an intermediate DeepSeek V4.1 Flash build. DeepSeek says the version uses a new model architecture, is natively multimodal, and is stronger, faster, and lower-cost.
Keep the existing base_url and set the model name to deepseek-v4.1-flash-expires-on-0910. Billing temporarily matches deepseek-v4-flash, with a limit of 20 concurrent requests per account.
NVIDIA (@nvidia) agreed to acquire Hugging Face for $12,930,300,000.
The first six digits, 129303, are Unicode code point U+1F917 — the 🤗 emoji. Written as the hex color #129303, they also make a green very close to NVIDIA’s signature color.
https://t.co/Y00BQsmLJh
OpenAI released GPT-6 Astra on Sept. 3, 2026, calling it a generational leap and saying it may mark the arrival of AGI. In a press briefing, co-founder Greg Brockman closed with “Welcome to the AGI era.”
OpenAI says Astra was its largest-scale training run yet — more than 100,000 GPUs at Stargate in Texas — and that other models played a major role supervising training for the first time.
Access starts with limited Daybreak Access institutions, then rolls out over the coming days to ChatGPT Plus, Pro, Business and Enterprise, the API (gpt-6-astra), AWS Bedrock and Microsoft Azure. API Standard is $10 / $50 per million input/output tokens; Fast is up to about 2.5× Standard speed at 2× the price. OpenAI had already named Astra the first Critical cybersecurity model under its Preparedness Framework.
Reportedly, OpenAI may move up its Astra model launch to Sept. 3, 2026.
A day earlier, OpenAI said Astra is the first model to meet the Critical cybersecurity capability threshold under its Preparedness Framework — meaning, with the right tools and access, it can find previously unknown flaws and develop exploits across many well-protected systems without step-by-step human guidance. The company still frames public availability as "soon," with advanced cyber access more limited.
https://t.co/hMfICNSzaV
Google (@GoogleDeepMind) released Gemini 3.8 Flash on Sept. 2, 2026. Built on Gemini 3.7 Flash, it targets software engineering and agentic knowledge work, with adjustable effort. It accepts text, image, audio and video, with a 1 million-token context and up to 64K output tokens.
On Google's official comparison table, DeepSWE v1.1 rises from 65.3% on 3.7 Flash to 71.0%, and Terminal-Bench 2.1 from 85.8% to 89.4%. Vals Finance Agent v2 (61.4%), Harvey's Legal Agent all-pass rate (10.0%), and HLE-Verified (54.9%) lead the table. Versus Claude Opus 5 it can match or beat several finance, legal, long-video, and Terminal-Bench 2.1 scores, but still trails on Terminal-Bench 4.0 (19.1% vs. 51.8%) and OSWorld-2.0 (59.0% vs. 75.4%).
API pricing matches 3.7 Flash's introductory rates: $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, 2026, then $1.50 / $7.50 from Jan. 1, 2027.
https://t.co/cfQJv400FO
Anthropic (@AnthropicAI) released Claude Fable 5.1 and Claude Mythos 5.1 on Sept. 2, 2026. They are the same underlying model with different safeguards: Fable 5.1 is generally available; Mythos 5.1 is limited to trusted access for cybersecurity and life sciences work.
Both have a 1 million-token context and up to 128K output tokens. API pricing stays at $10 per million input tokens and $50 per million output tokens, while cache reads drop to $0.25 per million tokens — 75% below Fable 5. Anthropic says that cuts typical workloads by about 25%, and highly agentic tasks by up to about 45%.
https://t.co/kjvnRcV4iy
Tencent Hunyuan (@TencentHunyuan) open-sourced Hy4 preview on Aug. 28, 2026: a mixture-of-experts (MoE) model with 770B total parameters, 49B activated per token, and a 1 million-token context. Tencent says it sits at the open-source frontier for coding, office, and science work.
https://t.co/SgiC7di0QU
https://t.co/jzHZmpAQlx
https://t.co/rB7ug5x9xp (@Zai_org) released GLM-5.3-Flash on Aug. 26, 2026, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B activated per token. On Artificial Analysis Intelligence Index v4.1.1 it scores 57 at $0.045 per task (discounted).
https://t.co/rB7ug5x9xp says it outperforms GLM-5.2 and approaches Claude Opus 4.8 on coding and agentic benchmarks. On in-house https://t.co/rB7ug5x9xp Code Bench v1.0 at max effort it scored 29.0 versus 29.5 for Opus 4.8. Before launch it ran anonymously as ox-alpha on OpenCode and OpenRouter, becoming the week's most popular model, with that traffic served on Chinese AI chips.
https://t.co/Y5DaUu87Ov
https://t.co/oMb9pDr4fA
Alibaba's Qwen team (@Alibaba_Qwen) released Qwen3.8-Flash-Next on Aug. 26, 2026: an open-weight multimodal mixture-of-experts (MoE) model with 125B total parameters and 6B activated per token. Qwen calls it an experimental preview of the Qwen4 architecture.
Every model in the world, at a glance.
CCHP 独家维护的云端价格表现已全新升级,带给你更丰富准确的模型数据,更快的更新频率,以及一个更实用的模型广场。
前往 CCHP 官网,开启全新体验⬇️
https://t.co/Tsp10RAJVq
*new-api 兼容的价格表格式即将支持。
Just wanted to say a huge thanks to @OpenAI for supporting CCH with 6 months of ChatGPT Pro!
I'll be trying to use Codex more frequently to build CCHP now—hopefully, this speeds up the release of CCHP! 🚀