DeepSeek V4 Pro GA estimates:
Terminal Bench 2.1 → 88-92
DeepSWE → 55-65
Cybergym → 78-85
Based on the exact jump Flash just made with post-training. Pro has 3.8× more active parameters.
This might actually challenge GPT-5.6 Sol. The open-weight king is loading…
Telegram just got pulled from the Apple App Store worldwide 🚨
No official reason yet. Existing installs still work. New iOS downloads are blocked.
For anyone running agents through Telegram (Hermes and similar harnesses):
The agent itself is fine.
Telegram is just the delivery channel.
If you already have the app → keep using it.
If you don’t → web version, desktop, or switch the gateway.
This is a good reminder: never make a single messaging app a single point of failure for production agents.
Build the system so the harness survives the client.
Qwen3.8 Max improved its agentic scores, but the operational picture is mixed.
Artificial Analysis measured it at $1.76 per task; roughly 2× Kimi K3 and 3× GLM-5.2. It also used 50% more output tokens than its predecessor, while hallucinations increased on one evaluation.
I haven’t tested it, so I don’t have a verdict.
Benchmarks earn a model a place in the evaluation queue, not in production.