GPT-6 Astra vs Fable 5.1 play chess.
When both models are set to "medium effort" with fast mode enabled for codex, the opposite happens.
GPT-6 processes each turn faster than Fable 5.1 compared to when both models are on max effort.
Which leads to GPT-6 winning on the basis of Fable 5.1 running out of time.
GPT-6 Astra vs. Fable 5.1 playing chess.
On max effort, GPT-6 takes consumes more time per move compared to Fable 5.1. Leading it to win on the basis of time.
Our previous tests with Qwen3.8-flash alongside other flash models in robotic manipulation where each llm must rotate the Rubik's cube once.
Perhaps with greater multimodal capabilities Qwen3.8-Omni-Flash can perform better on this task.
π Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities!
Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result.
Highlights: π₯³
- Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps.
- A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench.
- 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench.
Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever.
To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! π οΈ
We can't wait to see what you build with Qwen3.8-Omni-Flash! π
- Blog: https://t.co/oM9V1TkqYF
- Qwencloud: https://t.co/Cc2I8ELAnD
- Qwen Studio: https://t.co/V7RmqMaVNZ
- API: https://t.co/lNE7fH5YUt
- Qwen-MM-Plugins: https://t.co/SnM27dDP3d
- Qwen-Live Harness: coming soon
https://t.co/iVYlGjIbdy
Did GPT-6 Astra Max get nerfed?
Many have noticed a decline in performance from Astra recently. So we put it to the test, by comparing it to our launch day benchmark test.
Same Cathedral prompt. Same model and Max effort.
Launch: 26m 21.4s, $10.78
Today: 24m 34.7s, $9.85
Tested Claude Opus 5 across nine frozen graphics and game-development tasks against Claude Fable 5, GPT-5.6 Sol and Kimi K3.
Each run received one request and used its providerβs disclosed agent stack.
I gave Claude Opus 5 one prompt and let it work in Unity for 7 hours and 41 minutes.
It ended up with this walkable Japanese shrine.
Everything from the samurai, animations, torii, buildings, lanterns, forest, lighting and shaders was built by Claude.
Claude Opus 5 Max vs Claude Fable 5 Max vs GPT 5.6 Sol Ultra vs Kimi K3 Max make a cathedral from scratch.
All models received the same prompt, one shot.
Cost:
Claude Opus 5 Max: $29.25
Claude Fable 5 Max: $43.85
GPT 5.6 Sol Ultra: $5.61
Kimi K3 Max: $0.92
Wall clock:
Claude Opus 5 Max: 1h 41m 22.5s
Claude Fable 5 Max: 31m 13.8s
GPT 5.6 Sol Ultra: 38m 22.4s
Kimi K3 Max: 21m 54.1s
Claude Opus 5 Max vs Claude Fable 5 Max vs GPT 5.6 Sol Ultra vs Kimi K3 Max on a Neo-gothic city prompt inspired by @emollick.
All models received the same prompt, one shot.
Cost:
Claude Opus 5 Max: $19.55
Claude Fable 5 Max: $11.39
GPT 5.6 Sol Ultra: $2.84
Kimi K3 Max: $2.30
Wall clock:
Claude Opus 5 Max: 1h 52m 04.4s
Claude Fable 5 Max: 28m 48s
GPT 5.6 Sol Ultra: 22m 10.9s
Kimi K3 Max: 38m 29.1s
Claude Opus 5 Max vs Claude Fable 5 Max vs GPT 5.6 Sol Ultra vs Kimi K3 Max make No Manβs Sky from scratch.
All models received the same prompt, one shot.
Cost:
Claude Opus 5 Max: $133.83
Claude Fable 5 Max: $280.14
GPT 5.6 Sol Ultra: $20.84
Kimi K3 Max: $15.58
Wall clock:
Claude Opus 5 Max: 2h 31m 33.5s
Claude Fable 5 Max: 2h 05m 20.4s
GPT 5.6 Sol Ultra: 1h 29m 57.7s
Kimi K3 Max: 2h 23m 18.1s
Claude Opus 5 max vs Claude Fable 5 max make No Mans Sky from scratch.
Both models received the same prompt, one shot.
Cost:
Claude Opus 5 max: $133.83
Claude Fable 5 max: $280.14
Claude Opus 5 max vs Kimi K3 max on a prompt inspired by @chrisGPT. Both models received the same prompt, one shot.
Cost:
Claude Opus 5 max: $124.35
Kimi K3 max: $28.72
We tested Claude Opus 5 max against GPT 5.6 Sol max on a prompt inspired by @chrisGPT. Both models received the same prompt, one shot.
Cost:
Claude Opus 5 max: $124.35
GPT-5.6 sol max: $19.10
We tested Qwen 3.8 max preview against three other models. All models received the same prompt, one shot.
Cost:
Fable 5 max: $11.17
GPT-5.6 sol ultra: $6.86
Kimi K3 max: $0.45
Qwen 3.8 max preview: $0.30
Which model would you say performed the best?