Crazy. Opus 5 outperformed Fable 5 on 3d modeling at 2/3 the price
We tested Opus 5, Fable 5, and GPT 5.6 Sol on the same prompt: build an explorable Japanese suburban street in Three.js, full 3D, rendered like a hand-drawn anime background. All on Max effort.
Opus 5 took the longest but delivered the most varied scenery: clean dark outlines, flat cel-shaded color bands, dense street details down to the vending machine signage.
Fable paid more attention to details, and while GPT 5.6 Sol was 3-4x cheaper and the fastest, but the render shows it: sparse, washed out, barely styled.
Opus 5: 1hr 21min Β· $26.7
Fable 5: 1hr 5min Β· $35
GPT 5.6 Sol: 41min Β· $7.1
We don't just provide models randomly. All of these tokens were chosen with the expectation that everyone will circle back to using the best list of models.
So, here we go!
DeepSeek V4 Pro is now 61% off on GMI for unlimited time
$0.678 per 1M token input
$1.357 per 1M token output
available on open router and via GMI api key π
My predictions for 2026:
Coding and Mathematics AGI
- METR 50% time horizons above 24 hours - my mean estimate is 30.8 hours, 2 day time horizons possible within frontier labs when accounting for 60 day lag
- if 2025 was the year of agents, then 2026 will be the year of multi-agent systems
- agents delegating work to subagents -> the start of the agent economy and the great unhobbling!
Most of our current math and coding benchmarks will get saturated!
- Epoch Capabilities Index ( > 175 )
- FrontierMath Levels 1-3 ( > 95% )
- ARC-AGI 1 and 2 ( > 95% )
- SimpleQA verified ( > 95% )
- Simple-Bench ( > 90% )
- SWE-Bench-verified ( > 90% )
- Terminal-Bench 2 ( > 90% )
- WeirdML v2 ( > 85% )
- Humanities Last Exam ( > 80% )
- FrontierMath Level 4 ( > 75% )
- Cybench ( > 70% )
- GDPval ( > 70 % win rate, no ties)
- GSO ( > 65% )
- ARC-AGI-3 ( > 60% and > 80% if they go for o3-preview comparable compute budgets or continual learning breakthrough happens)
- more evals like gdpval that capture economic value of models and systems
- big focus white collar work and large acceleration of science: specifically i see acceleration in medicine, biology, chemistry, finance, legal, administrative work
- automation of white collar work will be enabled by having reliable and fast computer use agents
- reliable computer use agents will also have implications for how you use the internet. this is OpenAI's big goal: become the hub to the internet and delegate shopping and whatever to agents!
Big models launches to get hyped for in 2026:
- Claude 5 - Claude 5.5
- Gemini 3.5 - Gemini-4
- GPT-5.3 - GPT-6
(everything in between possible, but Gemini 4 ~ 80%, Claude 5.5 ~ 70%, GPT-6 ~ 60% likely before 2027)
- DeepSeek-V4
- Grok-5
- Qwen-4
- Kimi-K3, GLM-5, MiniMax M3
- more korean models and a bunch of american open-source models :)
The gap between closed and open labs will narrow in H1 2026 due to DeepSeek-V4, then widen in the later half of the year, especially on economically valuable tasks.
Closed models will be much more reliable. But we will still have Opus 4.5+ level open models by the end of 2026.
Most frontier models will be around 5-10T params. If we see GPT-6 and Gemini-4 at the end of 2026 10T+ param models are possible. These models + harnesses will be the first not research agents. We should also see much better live models with voice and video mode.
Model architecture:
- we will see both, more efficient architectures and more expressive architectures!
- hybrid architectures for even longer context windows, diffusion models for speed on edge devices, but also models that double down on full attention or even more expressive attention mechanisms
- looped language models, other recurrent architectures and continual learning will enable much smaller reasoning models! (TRM on ARC-AGI has paved the way for the reasoning core)
- big improvements in reasoning efficiency
in my 2025 prediction I included a prediction for 2026 that I stand by:
- "someone (Anthropic) figures out efficient test-time-training [...], this will be the next paradigm for 2026 and lead to superintelligence"
General outlook and some random thoughts:
- it will be clear to everybody that Anthropic has the mandate and is ahead of everyone else
- OpenAI, Anthropic and Google will remain frontier labs
- decent chance that Anthropic overtakes OpenAI's valuation and both are valued > 1T
- DeepSeek will join them with V4 as THE chinese frontier lab
- xAI will likely repeat Grok-4, Grok-5 will be great on benchmarks but Elon persists on slop-maxxing the model
- AI generated video content will take off with Veo-4 and Sora-3, consistent minute long videos will be possible
- embodied intelligence will start to take off by RL through world models
- full self-driving solved, waymo and tesla everywhere
- the stock market will have a 20%+ drawdown
- 15% chance of OpenAI going bankrupt and getting acquired by Microsoft due to collapse of oracle or a market crash, caused by rapidly deteriorating economic situation (unemployment, inflation)
- push against AI will become a common theme in most advanced western economies as unemployment rises
- populist right wing parties continue to gain traction in europe
- trump/republicans will lose midterm elections
@mathisdittrich This is awesome. We love browser use and would love to sponsor K3 and other OSS credits from GMIcloud to build more cool new agent use cases with the team there. Whatβs your thought.
Heyy there! GMIβs inference API is OpenAI-compatible, so any VS Code extension that lets you set a custom base URL + API key works. Point the extension at https://t.co/WqWte9hf2J with a GMI key and youβre running GMI models inside VS Code. That covers the popular coding-assistant extensions (Continue, Cline, Roo Code, etc.), and GMIβs own docs already call out Claude Code, Codex, and Cursor as supported clients.
Exciting update for our community
Today we announced:
- 2.4x live ARR growth in H1 2026
- $500M+ contracted ARR, 8x our year-end 2025 ARR
- Inference up 34x in 4 weeks, now ~2.5 trillion tokens/week
More capacity, more regions, more models coming online this year.
Thanks for building with us π
Read the full announcement: https://t.co/2z93vsY1lU
Qwen 3.8 Preview is quietly catching Kimi K3 on frontend + game design.And this is just the preview. Full 3.8 is going to be a beast.
We tested Qwen 3.8 Preview against Kimi K3 on 3D modeling and animation. Ran both through the same 3 prompts. Both on Max effort.
Qwen 3.8 Preview was dramatically faster across the board, 3-4x quicker on every test. Kimi K3 needed 2 hours on the final scene; Qwen wrapped it in 31 minutes.
Test 1 (Engine Sim)
Qwen 3.8 Preview: 25min
Kimi K3: 1hr 30min
Test 2 (Voxel Dino Printer)
Qwen 3.8 Preview: 8min
Kimi K3: 45min
Test 3 (Garden Scene)
Qwen 3.8 Preview: 31min
Kimi K3: 2hrs
I tried this new thing last week.
My hermes agents now have a daily standup with eachother.
I'm not involved.
It's fascinating and also a really good way to keep everyone in sync.
I'll be making a video soon on some of the ways I'm fasciliatating this.
meet Dino, our developer representative
It has a chip body cuz compute is the foundation for everything we develop here at GMI
we hope to bring our API key (or Dino) to more coding platforms in the future
let us know where you'd want to use GMI API key!
dino emojis on the way, and we'll bring it to our discord as well
Inking by @thinkymachines is now available to all GMI users
and we've made it super accessible at $0.467 per input, and $1.17 per output
check it out π
OFFICIALLY SAYING KIMI K3 IS BETTER THAN CLAUDE FABLE-5 AND GPT-5.6-SOL FOR FRONTEND DEVELOPMENT
Tried building a macOS-style UI for an AI operating system with Kimi K3.
This whole thing took about 15 minutes from a single prompt.
Inkling isn't the American GLM 5.2, it's more like an American Kimi K2.7 +
we tested Inkling, GLM 5.2, and Kimi K2.7 by having each build three games from scratch: a Tetris variant, a zombie tank survival game, and a 3D gladiator arena
on visual output, GLM 5.2 came out strongest, with Inkling in second
on speed, Inkling beat GLM 5.2 in all three tests, finishing in 12β18 min while GLM took 25β32 min every time
but on cost, Inkling was pricier than GLM in 2 of 3 tests, and K2.7 was the cheapest
Test 1: Inkling / $1.05 vs. GLM 5.2 / $0.70 vs. Kimi K2.7 / $0.10
Test 2: Inkling / $0.96 vs. GLM 5.2 / $0.79 vs. Kimi K2.7 / $0.39
Test 3: Inkling / $0.98 vs. GLM 5.2 / $1.10 vs. Kimi K2.7 / $0.22