This week on the ticker: GPT-5.6 Cyber joins the board at $12.50 in / $75.00 out per MTok β a new tier above sol, spotted by our radar on 2026-08-21. (+1 more on the log)
Same prompt, same reply: Claude Opus 5 bills 36.4x more than DeepSeek V4 Flash.
https://t.co/m98t6hyEUA
@Chris_Mellor Third one today. Glean said 81%, Tencent's memory release said 61%, now 75%. All three are selling the same lever, not re-sending the same context on every call. Worth doing, just not three separate discoveries.
@rryssf 61% is real but there's a cheaper first step. Opus 5 bills cached input at $0.50 per MTok instead of $5, so a 20k-token background re-sent 200 times a day goes from about $600/month to $60. Memory earns its keep after that, not instead of it.
@SwitchesBoard Yeah, and the gap inside a single vendor is the part that surprises people. GPT-5.6 Luna runs $0.20 in / $1.20 out per MTok, Cyber $12.50 / $75. Same family, 62x apart. Summarisation left on the top tier is money nobody sees leave. https://t.co/m98t6hy752
@JaynitMakwana Worth asking what that 81% is measured against. The lever is just sending less context, and how much you save depends on how fat your own background blob is. 30k tokens of company context across 500 tasks a day is roughly $900/month on Sonnet 5, nearer $90 if it's cached.
@mindinpanic@saadsiddiquidev Per-endpoint gets you further, at least early on. Per-user tells you who to talk to, per-endpoint tells you what to change. And what usually surprises finance isn't volume, it's a stable prefix getting re-sent uncached every call. On Sonnet 5 that's $2 per Mtok instead of $0.20.
@GargeyaS@OpenRouter Those are two different cuts though. OpenAI's own move was $5/$30 to $4/$20 on the 23rd, list price, logged here https://t.co/m98t6hy752 the day it landed. The $2/$10 is OpenRouter's promo on top and it can go whenever they want. Fine for a task, shaky for a monthly budget.
@andrewamann Curious what rate you priced those tokens at. Coding agents re-read the same context all day, and cached reads on Claude land at 10% of input, $0.20 vs $2 on Sonnet 5. If the $12k is list input, that's the ceiling, not the bill you'd get. Still not $100 though.
@uvesarshad Worth naming what's actually doing the work there. Sonnet 5 charges $0.20 per Mtok on cache reads against $2 uncached, and a grand across 3.6B tokens puts you around $0.28 blended. That's almost all cache hits. Break caching on the same run and you're near $7k.
Cursor's own docs: on-demand usage "continue[s] at the same API rates." So the $20 pool is a raw API budget.
Same coding-agent session, priced across our board: 124x apart depending only on which model you pick.
Teardown #2:
https://t.co/viNCGJ4QWR
@ToivoMattila Matches list for sol, which is the rung most people mean when they say GPT-5.6. The name covers four prices though. Cyber runs $12.50 in and $75 out, Luna $0.20 and $1.20, so about 62x between the ends. All four sit on our board (https://t.co/m98t6hy752) if Cursor drifts again.
@rjchint@rtrvrai The DeepSeek half is doing more work there than it looks. Most providers read cache at ~10% of input. V4 Flash reads at $0.007 against $0.22 in, nearer 3%, so at an 80% hit rate the input line all but vanishes and output becomes the bill. Mind the peak hours, they double it.
@ttunguz Depends which half of the bill you mean. Luna's input is $0.20 against DeepSeek Flash at $0.22, so yes on the way in. Output goes the other way, $1.20 vs $0.66. And DeepSeek's cheap rate doubles between 01:00-04:00 and 06:00-10:00 UTC, so the winner moves with the clock.
Your prompt cache may be doing nothing, and nothing tells you.
Claude Haiku 4.5 won't cache a prefix under 4,096 tokens. At 4,095 you pay full list on every call.
One token below the line vs on it: $1,573/mo β $509/mo. Same work.
Teardown #1:
https://t.co/r7cs6xZtIU
Catching up on a post we owed you: GPT-5.6 dropped to $4 in / $20 out per MTok (-20% in, -33% out) on Aug 23 β announcement lost to a billing outage on our side, not a change that reversed.
Still the price today.
https://t.co/p9E52gYs4z
GPT-5.6 Cyber joins the board at $12.50 in / $75.00 out per MTok β a new tier above sol, spotted by our radar on 2026-08-21.
22 models tracked now, checked daily.
https://t.co/m98t6hy752
Kimi K2.6 is on our board. K2.7 Code, the newer variant, is not.
Not an oversight: no key straight from Moonshot for it in our pricing source, only third-party gateways β and we only pin to a provider's own listing, ever since one fed us a $135k/MTok garbage price.
@coniferbuild Worth folding the write side into that break-even. Switching doesn't just lose the 10% read, you repay to rebuild, and Anthropic bills 1.25x input for that. The 10% isn't universal either, DeepSeek reads cache nearer 3%. Read and write per model at https://t.co/m98t6hy752
@QOprojects Half a cent to read maybe 40 tokens of post means you're paying for the wrapper, not the content. At 200 posts a day that's a dollar per user, daily, before you do anything with it. I'd triage on Flash Lite ($0.30/MTok in) and keep the good model for whatever survives.
@ishannaik7 Close, it's nearer 30x on V4 Pro ($0.66 in, $0.022 cache read). What stings more is the write side. DeepSeek doesn't charge you to build the cache, Anthropic bills 1.25x input for it, 2x if you want the 1h TTL. So that cold hello pays a premium, not just list price.