This framing is the right way to evaluate any AI spend: cost per hour of compute vs revenue per hour it generates, fully loaded. Most teams still budget AI like software (flat subscription, ignore utilization). Pricing it like labor with a P&L per agent changes every decision about what runs and when.
The gap I see between teams getting real value and teams stuck in pilots is never the model choice. It's whether they rebuilt the workflow around the model. Routing between providers is a commodity now. Encoding your domain's judgment into evals, guardrails, and context is the part nobody can buy off the shelf.
The industry is quietly running an experiment: what happens to the senior pipeline when nobody hires the juniors who would become them. NY Fed data already shows recent grads running higher unemployment than the workforce overall, first time in decades of records. Mentorship was never charity. It was how firms manufactured their own seniors.
AI made outbound free, so the scarce asset flipped from reach to attention filtering. The endgame is your agent answering your phone: it screens, verifies, wastes the spammer's compute, and only known humans get through. Whoever ships the personal gatekeeper that actually works owns the most valuable position in communication.
The elegant part is it prices conviction instead of talk. Every investor claims high conviction; almost none will concentrate a fund on it. Indexing off the GP's own position sizing turns cheap words into a costly signal. Founders should read term sheets the same way: ignore the enthusiasm, look at what percentage of their fund you are.
Outcome pricing works when the outcome is a dollar figure a third party confirms. We price denial appeals on recovered revenue: the payer's remittance is the referee, nobody argues about what counts. Tickets fail that test because the vendor grades its own homework. Pick outcomes with an external scoreboard and the pricing conversation gets easy.
Document parsing is the least glamorous bottleneck in applied AI and the most expensive when it fails. Healthcare still runs on faxes, scans of scans, and handwriting in margins. One misread field on a claim cascades into a denial weeks later. Benchmarks on clean PDFs flatter every model. Test on your ugliest documents first.
Simpler read: safety posts are enterprise marketing now. I sell AI into healthcare and every serious buyer asks about guardrails before they ask about capability. The labs learned the deal only closes when the CISO relaxes. Posting about safety converts better than posting about benchmarks, whether or not anyone got scared.
Code got CI before it got fast. Agents are getting fast before they get CI, and that ordering is the whole risk. Staging environments where an agent has to survive adversarial scenarios before touching production data should be as unremarkable as a test suite. Today they're a novelty worth a press release.
Agent skills right now are npm circa 2013: everyone installs, nobody audits, and the payload runs with your credentials. We started vetting third-party skills like we vet hires because a skill is an employee you never interviewed. Scanning before execution becoming table stakes was inevitable. The surprising part is how few people run any check at all.
The year of progress was mostly harness, and the model upgrades get the credit. Tests that run themselves, deploys that roll back, agents that check each other's work. Raw capability in August 2025 was already enough for working prototypes. The scaffolding around it was what turned demos into things that survive contact with users.
Local inference is quietly the biggest compliance unlock in regulated industries. Half my healthcare conversations stall on one question: where does the data go. A good-enough model running on hardware the clinic owns makes the answer "nowhere," and that closes deals cloud APIs can't touch.
Four companies spent $511 billion on capex in the past 12 months: Microsoft, Amazon, Google, and Meta. The buildout is led by AI data centers.
In 2021 the same four spent $125 billion. That's 4x in five years, straight from their SEC filings. It works out to $1.4 billion a day, over 40% of the entire US defense budget.
Every buildout this size in American history ran through the public: railroads got land grants, the grid got regulated, highways got acts of Congress. This one is four boards answering to shareholders.
Fastest infrastructure deployment ever, or the least accountable? Can it be both?
At 10,000 tokens a second the constraint moves entirely to the human side of the loop. Generation becomes instant, so review, approval, and deciding what to build become the whole job. Organizations built around waiting for work to finish have no idea what to do when the work finishes before the meeting ends.
The national numbers already say this quietly. Fed data: top 10% of households own 87% of the stock market, bottom 50% own 1%. SF compresses that whole distribution into ten square miles where both ends ride the same bus line. The city renders the national gap at full resolution, and AI money is turning up the contrast.
Open sourcing the loop infrastructure is the smart move: the loop was never the moat. The moat is the proprietary signal feeding it, and Shopify sits on one of the best commerce signal streams in the world. Giving away the engine while owning the fuel is how you set a standard and keep the advantage.
The capability line more than doubled its slope. The adoption line didn't move: hiring still frozen, most enterprises still piloting, workflows still designed for humans. That widening gap between what models can do and what organizations absorb is where the next decade of companies gets built.
Verification is the bottleneck everywhere, math is just where it's most visible. In production agent systems, generation stopped being the constraint a year ago. The scarce asset is the harness that proves an output is correct cheaply. Whoever industrializes verification for a domain quietly captures that domain's margin.
The same bet structure works one layer down: build the product that's barely economical at today's token prices and let the cost curve do the heavy lifting. Inference cost per unit of capability has been falling faster than launch costs ever did. Companies priced for today's costs are leaving the compounding on the table.
Selling AI into healthcare made this brutally concrete for me. The cost of the sale is trust-building, and trust costs the same whether the contract is $50K or $500K: same security review, same pilot, same skeptics in the room. The fix is pricing on risk absorbed and outcomes owned, then refusing to run the trust gauntlet for small checks.