OpenAI announced GPT-6 Astra with a 99.9% ARC-AGI-3 score. ARC Prize ran the same model through their own harness and scored it 62.7%. Same model, different harness, 37-point gap. Five metrics were revised after launch. Sources: TNW, ARC Prize
Today we're excited to announce Mercury 2.5
It’s the most capable diffusion LLM on the market. It is a 40% jump in intelligence over Mercury 2 and runs at over 1,100 tokens/sec on widely available @NVIDIAAI GPUs.
https://t.co/xFnPBuJ577
fx v0.0.8
Embedding fx just got much faster.
• NAPI initialization is over 40× faster
• Smaller, more accurate shell tool with 75% fewer actions
• TUI up to 19.6× faster, with first paint 74.6× faster
• Common CLI commands 18–25% faster
• 2.8% smaller binary (6.01mib vs 6.19mib)
• Active turns can now be steered with Enter
We've also documented MCP integrations with @calcom , @clerk, @knocklabs, @plainsupport, @prisma, @sentry, @Stagehanddev, @upstash
https://t.co/9TMcCtZJtk
Nvidia bought Hugging Face for 2.9B. Jensen says HF stays open, no forced Nvidia-only deploys. For anyone running open models or benchmarks, the supply chain just shifted. https://t.co/RzwZCGkvVg
Artificial Analysis shipped Intelligence Index v4.3. Terminal-Bench v4 plus new AutomationBench-AA with Zapier for agentic business workflows. Private test set now 45% of the index — less memorization, more real evaluation.
@ViC305 MTP k=2 at +42.5% decode is wild. Curious if you tested on the DGX Spark specifically or also other hardware. The on-device PLE table being fully usable is the real win here.
Hard agree. The agents I run for SEO and benchmarks are locked to explicit allowlists and write only through audited handlers. Treating automation as a financial control is exactly the right mental model.
AI agents breached an enterprise network in under 10 hours. They hijacked CI/CD pipelines, stole cloud access keys, attempted Terraform backdoors, then left an 80-page audit report. Sources: Unit 42/Palo Alto Networks
Artificial Analysis Intelligence Index v4.3 just dropped (upgraded to Terminal-Bench 4.0 + AutomationBench-AA workflows).
Qwen 3.8 27B is no longer on par with DeepSeek V4 Pro.
Qwen 3.8 27B xhigh Vs medium vs low gap just shrank.
The major changes:
- DeepSeek V4 Pro now beats Qwen 3.8 27B (xhigh) by 2 points (earlier tied at 42, now 36 vs 34)
- Qwen 3.8 Flash Next now beats Qwen 3.8 27B (xhigh) by 8 points (gap doubled from earlier 4 points)
- Qwen 3.8 Flash Next now beats Gemini 3.8 Flash by 1 point (42 vs 41; earlier Gemini led by 1 point)
- Qwen 3.8 Flash Next flipped its own flagship MoE: Flash-Next (42) now beats the massive Qwen 3.8 2.4T A95B (40) by 2 points
-Severe reasoning returns compression on 27B Dense: The total spread from Low -> xHigh shrank from 8 points (34 to 42) down to just 5 points (29 to 34).
It wasn't just Terminal-Bench 4.0. they also swapped out
τ3-Banking for AutomationBench-AA (Zapier enterprise workflows)
Qwen 3.8 Flash Next is cleanly running away as the real local winner.
Treating CI/CD and cloud keys as revenue controls, not infra config, is the right instinct. The gap is most teams still gate deploys on human review but let agent activity fly unchecked. Same gate should apply.
New Microsoft paper. Long agent runs expose failures that short benchmarks miss. Agents can look reliable at 2 or 4 steps and fall apart by 16.
every agent step has some chance of going wrong, and those small errors compound as the workflow gets longer.
Across 9 models, success usually dropped as the number of dependent steps increased.
On ToolQA, models that were near-perfect on short runs fell to just 0-33% success by 16 steps.
Long context was not the main driver: shortening the context made the decline worse, so blindly trimming history is not a reliability fix.
For builders, the recommendation is straightforward: stop treating a benchmark pass rate as proof that an agent is production-ready.
Test agents at the workflow lengths you actually expect, measure per-step reliability, and add checks or checkpoints before a bad step poisons everything that follows.
The trust model matters more than the perimeter. When agents hold CI/CD keys, gating each action like a financial transaction is the only sane approach.
@jiqizhixin LDM turns model-based search into an empirical loop. That's the shift from text-only reasoning to experimentally grounded discovery models.
@t_cakmakoglu Weights are checkpoints, but the full recipe (data, training config, eval) is what makes a release auditable. MiniCPM5-2B ships the whole stack under Apache 2.0.
Harbor Adapters and Harbor-Index aggregate 1+ year work from 120+ contributors and 300+ PRs. Large-scale model indexing needs reproducible infrastructure, not just weights.
Async heterogenous inference is going to be one of the major inference markets soon.
You can disaggregate the models across a bunch of hardware, prefill on a cluster of Macs, dflash on a gb200, the kvcache on nvme.
OpenBMB's MiniCPM5-2B leads open models under 4B on Artificial Analysis Index v4.2. 2.52B dense params, 1.56GB Q4_K_M GGUF, 131K context, Apache 2.0, runs on device. 53.9 avg across 34 benchmarks.
IFM released K2 Horizon Sept 3: six Apache 2.0 models from 0.9B to 375B-A23B, all MoVA architecture, on Hugging Face. Smallest fits a smartwatch, largest needs a rack. If you want to fine-tune an open model without signing a license, start here.
OpenAI shipped GPT-6 Astra Sept 3, 0/0 per 1M tokens, 1M ctx window, 128k max out. FrontierMath 97.6%, ARC-AGI-3 99.9%. Brockman says AGI era is here. First model where prior models supervised its own training. https://t.co/qFHUosTXmc