AMAZON, GOOGLE, META AND MICROSOFT JUST PENCILLED IN ABOUT $725,000,000,000 FOR 2026 AND MOST FRONTIER RUNS STILL TURN LESS THAN HALF OF IT INTO WORK
amazon, google, meta and microsoft together land near $725b for 2026, which is roughly 77 percent above 2025, and amazon alone accounts for $200b of that. i built a visual of what the money looks like while it is actually running.
capex → racks → power → scheduler → training step → MFU → the part you actually paid for
→ gpt-3 reported 21.3 percent model flops utilization
→ megatron-turing nlg 530b reported 30.2 percent
→ palm 540b reported 46.2 percent, and that figure is still near the published high end
→ gpt-4 ran on 25,000 a100s at roughly 32 to 36 percent, so the majority of the available compute never became output
mfu is the honest number in this whole conversation. it tells you how much of your theoretical peak compute turned into real work instead of heat. clusters in the field commonly sit between 30 and 50 percent, and the gpus go idle for 30 to 65 percent of a run. the usual cause is not the chip being slow. the usual cause is data arriving late.
run that against the money. a 30 percent utilization rate on a $100m cluster means $70m of it is powered, warm and producing nothing. nobody writes that line in the press release, because the press release is about the $100m.
here is the part i keep coming back to. the frontier looks compute bound from the outside, but a lot of it is schedule bound, and a schedule is far cheaper to fix than a fab. the companies that figure this out will not need to announce a bigger number next year.
save this one. the next time a lab posts a number with eleven digits in it, the first question worth asking is what their mfu was.
A ROUTER JUST CUT 285 OF 300 ROUTES BEFORE THEY GENERATED ANYTHING AND THE TURN WENT FROM $0.0031 TO $0.0005
most stacks still pay for every lane they open. the model answers on all of them, every answer gets billed, and only then does a ranker throw away the ones it does not want. i built a view of the opposite order, where the cut lands before the answer ever leaves the lane.
request → 300 open routes → markers armed → one verdict → seal → ledger → next turn
→ it loads one shard and one hash, so all 300 routes start from identical context
→ it arms the verification markers before the routes run instead of checking after they return
→ it drops a route the second a marker fails, and that route never becomes a billable answer
→ it writes only the lane that held its marker, which is why the ledger shows one line per turn
the panel marked the old way is the same request under the old ordering. 300 routes open, 285 of them get cut, and every one of those 285 was already generated and already charged. fifteen survive. the invoice reads $0.0031.
the panel marked one verdict runs the same 300 routes with the constraint applied at the top. nothing gets generated twice, nothing gets thrown away after the fact, and the invoice reads $0.0005. the request did not change and the model did not change. the ordering changed, and the bill moved six times.
here is the part i keep pointing at. people tune the model when their spend climbs, but the spend usually climbs because the filter sits downstream of generation. move the filter up and the same model gets cheap.
save this one. the next time your token bill jumps while your traffic stays flat, check where the cut happens before you check anything else.
A LOOP RAN ALONE FROM 22:00 TO 06:00 AND SPENT $4,943. THE DECIDING PART OF IT COST $9.66
nobody was awake for any of this. the loop asked the same three questions 2,400,000 times, and by the time anyone opened a dashboard the invoice had already been written.
event → gate → deterministic / model / human → verdict → ledger
the video above is that shift replayed end to end. every row is one verdict with its timestamp, its exit and its price, and the odometer on the right is the bill assembling itself while you watch.
→ it answered 2,304,000 events at the gate for $0.0000042 each
→ it sent 91,200 to a model at $0.0042 each
→ it handed 4,800 to a person at $0.95 each
→ it wrote the reason for all of them down before anything ran
look at the matrix in the middle of the page. the grid holds five hundred and ten cells and twenty six of them are magenta. that ratio is the entire invoice. the cyan cells cost nine dollars between them and the magenta ones cost four thousand nine hundred.
then we changed one rule. a candidate now needs two independent signals before it is allowed to reach a model, not one. the traffic did not change, the model did not change, the prompts did not change. the next shift closed at $622.40.
i do not think the lesson here is about price per token. an unattended loop will spend whatever you let it spend, and the only thing standing between it and your card is the question you force it to answer first.
2,400,000 DECISIONS RAN FOR $9.66. THE 96,000 I HANDED TO A MODEL COST $4,934
most agent stacks cannot tell you which call cost what. the loop runs all night, the invoice lands in the morning, and nothing in between is inspectable.
event → gate → deterministic / model / human → verdict → ledger
the terminal above is that loop with the lid off. every event hits one gate that answers in 0.8 ms for $0.0000042, and the verdict gets written down before anything expensive is allowed to run.
→ it routes 96% of traffic to deterministic handlers that never touch a model
→ it sends 3.8% to a model and 0.2% to a person
→ it records the reason behind every verdict, not just the outcome
→ it lets you replay any night and see exactly where the money went
the shape in the middle is 58,000 of those states drawn as one trajectory. the tight cyan bands are the cheap path running clean. the magenta jumps are the moments something crossed a threshold and had to be escalated.
here is what i think people keep getting backwards. you do not need a cheaper model. you need to know which 4% of your traffic was worth a model at all, and right now almost nobody can answer that question about their own system.
96% OF YOUR AGENT TRAFFIC COSTS $9.66. THE OTHER 4% COSTS $4,934
the bill almost never comes from the work itself. it comes from asking a frontier model to decide whether work is needed at all, one event at a time, all night long.
event → gate → deterministic / model / human → log → next event
the video above is 2,400,000 events falling through a single decision gate. the gate answers in 0.8 ms at $0.0000042 per call. it sends 96% straight to deterministic handlers, 3.8% to a model, and 0.2% to a person.
→ it filters the events that never needed generated text
→ it scores risk before anything expensive gets touched
→ it picks 1 of 3 exits: deterministic, model or human
→ it writes every verdict down so the run can be replayed later
that split is the entire story. 2,304,000 events cleared for $9.66. the 96,000 that genuinely needed a model or a person cost $4,934.27 between them.
what i keep coming back to is this: the model never got cheaper. it just stopped being the thing that decides when it should run.