This completely changes how I look at AI benchmarks.
Qwen 3.8 Max Preview felt nearly unusable in OpenCode and Qwen Code.
With the Claude Code harness, it suddenly came surprisingly close to Kimi K3 and Fable 5.
Same model. Same task. Different harness. Wildly different result.
How is this even possible?
Wow, if you connect Fable to image-gen and image-to-3D APIs, it's *insane* for game dev.
It made an infinite explorable universe, with characters and lore. All assets and sound completely from scratch.
3 prompts in 2 hours, max effort. Workflow below.
@xuanmingzhangai Don't listen to people asking for small models. Look at Anthropic revenue, people want large token efficient models an are willing to pay for it
@ChujieZheng I like this model. Here's Christmas celebration scene I told it to build. Probably not in the training data, so not benchmaxxed. https://t.co/su3nzeDkPP
@teortaxesTex You should try new V4 flash on the web instead of shit posting. It's actually good. I'm having fun testing it because it's fast and actually can improve mid results when you tell it to.
If you’ve tried Qwen3.8 Max Preview on OpenCode note that the effort level is the lowest by default
You have to configure this beforehand in OpenCode’s settings (just ask claude or codex to do it)
The efforts levels are none, minimal, low, medium, high and xhigh
xhigh is significantly better and yields amazing results