Inference got a hundred times cheaper this year. The compute bill went up anyway.
If you understand why those two sentences are both true at the same time, you understand the most important thing happening in AI right now.
I work on inference for a living, at @nebiustf, where we run open-source managed inference at scale. Most of what follows is what I'm seeing from inside the bill.
12 months ago, the cost of 1M tokens of frontier-class reasoning was somewhere on the order of $60.
Today, an equivalent quality of output costs roughly $0.50.
Price /token of o1-level intelligence has dropped about a 128x in a year.
Price of GPT-4-level output has dropped roughly 100x since the original GPT-4 shipped.
By any normal reading of a technology cost curve, this should be deflationary. It should be saving customers money.
The opposite has happened. The total compute bill at every hyperscaler is going up, not down. Anthropic just signed multi-year capacity deals with both XAI and Amazon. Microsoft's Azure capex guide for 2026 starts with an eight. OpenAI is reportedly spending more on compute every quarter than it did in all of 2023. Nvidia paid roughly twenty billion dollars to acquire Groq, an inference-specialist company that did not exist as a serious commercial entity three years ago.
The cost curve and the demand curve crossed, and then the demand curve lapped the cost curve.
Here is what happened underneath.
A reasoning model burns roughly 10x the output tokens of a non-reasoning model on the same task, because it spends most of its tokens thinking out loud before answering. An agentic workflow chains roughly twenty times the requests of a single-shot completion, because it loops, calls tools, plans, retries, and synthesizes. A modern deep-research query (the kind a research analyst can fire off in fifteen seconds and then walk away from for ten minutes) costs more compute than 10 original GPT-4 queries combined. We made every individual token a hundred times cheaper, and then we built a generation of products that consume ten thousand times more tokens.
This is the Jevons paradox playing out at trillion-dollar scale, in compressed time, in front of everyone. Jevons noticed in 1865 that making coal-burning more efficient did not reduce coal consumption. It increased it, because efficiency unlocked uses that were previously uneconomic. Steam engines became more practical at smaller scales. Whole industries that could not afford coal at the old price suddenly could. Britain's coal consumption rose sharply, not despite the efficiency gains, but because of them.
The same thing is happening to AI compute right now and it is happening faster than any analogous historical cycle. Falling token prices did not contract demand. They unlocked agents, deep research, code-writing systems, multi-step reasoning, persistent memory, the entire next layer of AI products. Every product in that next layer consumes orders of magnitude more compute than the chat interfaces it is replacing.
The math at the aggregate level is brutal: 100x cheaper tokens times 10 000 more tokens equals a 100x larger total bill.
The implications stack quickly.
If you are running a hyperscaler, your 2026 capex guide is not a peak. It is a step on a curve. Inference is structurally always-on, twenty-four hours a day, in a way that training never was. Training is bursty. You spin up a cluster, run for weeks or months, and stop. Inference runs continuously, scales with usage, and the usage curve is exponential. Your power bill, your cooling bill, your transceiver count, your storage footprint, all of these were sized for a workload mix that no longer exists.
If you are running an AI software company built on top of someone else's closed API, you have a problem that did not exist a year ago. Your gross margins get worse as your customers get more value out of your product, because the more they use it, the more compute you pay for. The companies that win this are the ones that figured out vertical integration before the math caught them.
If you are watching this from a distance and trying to understand where the next bottlenecks form, the answer is everywhere downstream of "more inference compute, always-on, with massive memory state per session." The KV cache, the running memory state of a long conversation or an agent loop, is the silent monster of the inference era. It does not scale linearly with parameters. It scales linearly with context length and number of agent steps. A long agent session can hold tens of gigabytes of state per user, per session.
Multiply that by every concurrent user of every product, and you understand why $MU, $SNDK, $TOWCF, and the entire memory and packaging layer have re-rated the way they have.
The CPU-to-GPU ratio is evolving. Training is 1:8. Basic chat inference is 1:4. Agentic inference is 1:1, sometimes CPU-heavy. Google has split its TPU line in two, with a dedicated inference chip carrying tripled SRAM for KV cache. $INTC and $AMD just spent two earnings calls explaining that this shift is structural, not cyclical. The hardware map is redrawing in real time and the financial press is mostly still writing about training clusters.
The right framing of where we are right now is not that AI is hitting a wall. The framing a year ago that scaling was hitting a wall was the most expensive bad take of the cycle. The right framing is that AI got dramatically cheaper, dramatically more capable, and dramatically more useful, and the cost of running it at the new equilibrium of demand is much higher than the cost at the old equilibrium of demand, because the new equilibrium is enormous.
A meaningful share of what we actually do at Token Factory, day to day, is help customers stop their bills from running away from them. KV-cache management. Speculative decoding. Quantization. Routing. The kind of vertical integration that, eighteen months ago, every product team was happy to leave abstracted away behind a closed API. The reason this stack matters now is the same reason this whole essay matters: at the new equilibrium of inference demand, the cost of treating compute as a commodity is no longer survivable. The companies that figure out the layer beneath the API are the ones who keep their margins.
Cheaper tokens. More tokens.
Same coal as 1865.
Sorry, but remember my cube concept?
X- Bitcoin price change changes the stocks price by the same amount assuming mNAV and btc/share stay the same.
Y- mNAV change affects price the same way, assuming Bitcoin price and btc/share stay the same.
Z- BTC/share change affects price as a simple ratio as well assuming Bitcoin price and mNAV stay the same.
All of this being said, graphs below are cute, but please rationalize them with not just the potential change in Bitcoin price. In order to perform differently relative to Bitcoin, it’s now the area of a 2D surface stretching or compressing with one axis being mNAV change from today (or your baseline) and the other axis being BTC/share change. So, if you expect mNAV to stay flat and BTC/share to 10x its current amount (impossible for Strategy who holds 3% of all Bitcoin and would require holding 30% with zero share dilution in order to 10x btc/share), but if that’s still the naive expectation, that would allow a max outperformance from here of 10x in Bitcoin terms.
“But what about stock buy backs or earning yield in fiat instead of Bitcoin?” As Bitcoin price appreciates (not talking 50% but 1,000%… 10k% and so on) and a companies stack gets larger, yield and the power of buybacks will shrink. Expected double digit % outperformance of Bitcoin will be a thing of the past in coming decades once geo markets are saturated. That’s why now is such a critical time for investors to research and see if anything is worthy of their conviction.
“But mNAV is uncapped if they just stop selling shares on common. Plus, it’s kinda like PE ratio”. Sorry, not uncapped. I disagree with lots of people in this space here, but I’ll try to explain why. Most companies don’t go around selling common shares to the market very often. So why is their price capped? Because they’re unfortunate enough to hold fiat instead of Bitcoin on their balance sheet? Give me a break 😅. Look, we all know Bitcoin will outperform fiat long term, but it is the Bitcoin per share yield which provides tangible outperformance of the underlying. You think MSTR can reach 5% of all Bitcoin (holding 1.05 million of those puppies?). Great, I do too. So, once they do, if their mNAV hit 20 they would have a market cap equivalent to the value of all 21 million bitcoin. If you still think that’s a rational and even likely outcome, consider why investors would bid it up so high. The goal and potential of MSTR and others to outperform Bitcoin has to be quantified in BTC/Share. Otherwise it’s similar to a Bitcoin ETF. Is IBIT trading at a premium because it holds 600k (or whatever amount) Bitcoin? Nope. Why? Sure, by design, but don’t fall back on it not being an operating company etc. the differentiator is yield. If Saylor stopped providing yield.. decided they have enough Bitcoin and they’ll just squat on it, you think mNAV of 1 isn’t the general result? We’ve seen mNAV below 1 (where they certainly weren’t selling new shares to the market).
“But Climb, that was during a bear market”. Ok, so every bull market MSTRs mNAV can go unchecked to 100 and we will just drop 90-99% or so during bear markets?
“But Climb, there won’t be any more bear markets”. Sirs and ladies, can humans still leverage and overextend themselves? Is margin trading still possible? Is individual collateralized debt still possible? Yes. Human greed makes pops and flops of volatile sectors a “when not if” phenomenon.
Want a short thread of things which will keep mNAV from hitting 100 for anything that has more than a handful of Bitcoin and is trading on hope with extremely low marketcap anyway? 🧵 👇
Execution
Extremely proud of the community and team for the smooth launch of the HyperEVM. The upgrade happened amidst billions of dollars of daily volume, where the majority of defi derivatives trade. There was no downtime, and no performance degradation after the launch.
The UX of trading is still so seamless that many users assume the HyperEVM is a separate chain! To be clear: the HyperEVM and the existing native Hyperliquid financial system are one composable state.
The safest way to upgrade this uniquely complex system is a gradual rollout. Precompiles and other L1 interactions will build upon the sturdy foundation of the initial HyperEVM release. This will unlock an entirely new class of performant defi applications, but more on that later.
Right now I want to focus on the execution of the launch. A bug in HyperEVM logic or an unoptimized code path would've crippled the entire blockchain, affecting hundreds of thousands of users and billions in open interest.
This was a massively challenging launch: a jet's engine was flawlessly changed mid-flight.
--
Philosophy
The HyperEVM launch stayed true to Hyperliquid’s “no insiders” principle. Hyperliquid has always embodied the original ethos of crypto: no investors, no paid market makers, no fees going to any company. The HyperEVM launch is yet another example that integrity and fairness are the pillars of Hyperliquid.
The tradeoff of a fair launch is that things are a bit messy to start. Tooling might not be there from day one. Builders need to familiarize themselves with the tech. But these short term obstacles are nothing compared to the long term value of fairness. No one had a head start or unfair advantages. I’m impressed that some teams deployed dapps, tooling, and analytics within hours of the HyperEVM release, a testament to the strength of the builders and community.
Hyperliquid will eventually be the credibly neutral infrastructure that houses all of finance. Looking back, the L1 launch, HYPE genesis, and HyperEVM launch will all be important milestones. Success is path dependent, and there can be no blemishes on the fair trajectory towards the final state.
The HyperEVM is a clean slate. The community is hungry for quality applications. Where else in the world is there such an imbalance between supply and demand for applications built? Fast, general purpose chains are nothing new. But on Hyperliquid, builders can plug into a mature, liquid, and performant onchain economy with real users.
I've noticed a pattern that builders, traders, and communities who "make it" on Hyperliquid are those who call Hyperliquid home. Legacy players don't win just because of their credentials. Newcomers have equal opportunity to win by challenging the status quo and seizing the opportunities. There are empires to be built on the HyperEVM, and the community welcomes builders with open arms.
Hyperliquid
“If you are frightened of debasement of your currency…you can have an internationally based instrument called Bitcoin….I’m a big believer…” - Larry Fink of @BlackRock