Benchmarks tell you which LLM is best. Usage tells you which model is winning.
Tidelines tracks the AI model market as it's actually used: market share, momentum, and prices using public gateway data such as @OpenRouter.
https://t.co/y3wYIGPvfM
Is routed inference getting cheaper?
The volume-weighted average price paid per token on OpenRouter is down 65% since mid July.
Over the same window, daily token volume and estimated daily spend on the gateway both hit records.
That average is our Token Price Index: list prices weighted by how much traffic each model gets. It moves when routing shifts, not when anyone changes their prices.
Three trends behind the decline:
1. The expensive models lost their traffic.
On 5 July, Anthropic's Opus and Sonnet handled about 11.6% of tokens and DeepSeek's V4 Pro another 6.2%. By 30 August none of them was above 2%. These are the priciest models in the basket, so their share leaving pulls the average down.
2. That traffic went to cheaper flash tiers.
https://t.co/nhBoBfYFlC's GLM 5.3 Flash is at 11.5%, DeepSeek's two V4 Flash releases at 10.8% and 5.1%, Tencent's Hy3 at 5.6%. That's about a third of all tokens in budget tiers that used to go to mid-range and premium models. Tencent's Hy4 preview is at 8.7% but doesn't belong here. It's roughly 6× the price of Hy3 and takes about 11% of dollar share.
3. Free variants gained share.
MiniMax's M3 free version is at 5.6%, NVIDIA's Nemotron 3 Ultra free version at 4.0%. Free tiers are a distribution play: give the tokens away, get developers routing to you, charge later. More labs have been running it this year.
Two limits on reading this. Tokens are not dollars, and on this same gateway Kimi K3 is about 1.7% of tokens and 13% of estimated spend. And a public router is one slice of demand, so nothing here tells you what a lab earns outside it.
What it is good for is a proxy: price-sensitive routed inference is getting cheaper, and Chinese labs plus NVIDIA's free Nemotron supply most of that volume.
The American labs already have partly adjusted to this. OpenAI's GPT-5.6 Luna is up 5.0pp to 10.7%, a cheap tier carrying a gateway discount on top of a list cut. NVIDIA is running the free-variant play. Google still holds Flash share.
Is routed inference getting cheaper?
The volume-weighted average price paid per token on OpenRouter is down 65% since mid July.
Over the same window, daily token volume and estimated daily spend on the gateway both hit records.
That average is our Token Price Index: list prices weighted by how much traffic each model gets. It moves when routing shifts, not when anyone changes their prices.
Three trends behind the decline:
1. The expensive models lost their traffic.
On 5 July, Anthropic's Opus and Sonnet handled about 11.6% of tokens and DeepSeek's V4 Pro another 6.2%. By 30 August none of them was above 2%. These are the priciest models in the basket, so their share leaving pulls the average down.
2. That traffic went to cheaper flash tiers.
https://t.co/nhBoBfYFlC's GLM 5.3 Flash is at 11.5%, DeepSeek's two V4 Flash releases at 10.8% and 5.1%, Tencent's Hy3 at 5.6%. That's about a third of all tokens in budget tiers that used to go to mid-range and premium models. Tencent's Hy4 preview is at 8.7% but doesn't belong here. It's roughly 6× the price of Hy3 and takes about 11% of dollar share.
3. Free variants gained share.
MiniMax's M3 free version is at 5.6%, NVIDIA's Nemotron 3 Ultra free version at 4.0%. Free tiers are a distribution play: give the tokens away, get developers routing to you, charge later. More labs have been running it this year.
Two limits on reading this. Tokens are not dollars, and on this same gateway Kimi K3 is about 1.7% of tokens and 13% of estimated spend. And a public router is one slice of demand, so nothing here tells you what a lab earns outside it.
What it is good for is a proxy: price-sensitive routed inference is getting cheaper, and Chinese labs plus NVIDIA's free Nemotron supply most of that volume.
The American labs already have partly adjusted to this. OpenAI's GPT-5.6 Luna is up 5.0pp to 10.7%, a cheap tier carrying a gateway discount on top of a list cut. NVIDIA is running the free-variant play. Google still holds Flash share.
What a session actually costs in Hermes @NousResearch per model.
Median cost of one coding session by session length. 30-day window, log scale.
Data via @OpenRouter
Recent usage growth of @OpenAI on OpenRouter is attributable to the latest models strategy.
GPT-5.6 Sol is the powerhouse flagship model optimized for deep, multi-step reasoning, complex programming.
GPT-5.6 Luna is the lightweight model optimized for high-volume, low-latency cost-efficient task execution.
So GPT 5.6 Luna takes most of
Token consumption, while GPT 5.6 Sol takes most of $ spend.
Recent usage growth of @OpenAI on OpenRouter is attributable to the latest models strategy.
GPT-5.6 Sol is the powerhouse flagship model optimized for deep, multi-step reasoning, complex programming.
GPT-5.6 Luna is the lightweight model optimized for high-volume, low-latency cost-efficient task execution.
So GPT 5.6 Luna takes most of
Token consumption, while GPT 5.6 Sol takes most of $ spend.
Hermes Agent (@NousResearch) is the #1 app by token consumption on @OpenRouter.
Looking at model usage mix on Hermes is interesting because it’s very easy to switch between models and users are normally on the look for the best model from a cost effectiveness perspective.
In the last days @deepseek_ai v4 Flash0731 raised to the top as the most used model.
Top models used on the last model were:
- @TencentHunyuan Hy3
- @deepseek_ai v4 Flash0423
- @StepFun_ai 3.7 Flash
- @NVIDIAAI Nemotron 3 Ultra
Token usage on @OpenRouter continues to grow, although there is a steep decline in users dollar spend per lab in the last two weeks.
This decline is mostly attributable to Anthropic models usage decline.
Two weeks ago 62% of the users spend happened on Anthropic models, that was reduced to 46% today.