Morgan Stanley: Token Costs
Data Center Net Margins From Token Sales
> High Overall Profitability: The intelligence factory model projects strong net margins ranging from 58% to 90% for token sales, heavily dependent on the type of GPU used.
> Feynman Data Center Leads: Feynman achieves the highest net margin, approaching 90%.
> Rubin Data Center: Demonstrates strong performance with a net margin of approximately 78%.
> Blackwell Data Center: Yields the lowest margin among the three generations, though still robust at around 59%.
Token Pricing Reduction by GPU Generation
> Significant Cost Reductions: Subsequent generations of NVIDIA GPUs drive down token pricing considerably compared to the baseline Blackwell architecture.
> Blackwell to Rubin: Transitioning from Blackwell to Rubin lowers token pricing by roughly 47%.
> Blackwell to Feynman: Transitioning from Blackwell to Feynman achieves a much steeper reduction, lowering token pricing by approximately 76%.
$NVDA
NAND shortage to ease into H2 2027?
According to TrendForce, the NAND flash supply growth is expected to outpace the demand growth in 2027, resulting in TrendForce expecting the shortage to ease by H2 2027.
While inference & agentic AI is expected to continue driving up demand for memory, the consumer electronics demand is expected to weaken in 2027. Combined with larger NAND supply, the shortage is expected to ease into H2 2027.
AI is increasingly becoming a larger share of total NAND demand (over 50% in 2027), but consumer electronics remain a large share too (over 35% in 2027).
Quite an interesting article if you ask me. It paints a much different picture than the news of NVIDIA $NVDA asking Samsung to increase its NAND production by the same amount as the demand for NAND from Apple $AAPL.
$SNDK $KXIAY $MU $DRAM $EWY
FT: CHINA’S MINISTRY OF COMMERCE HAS DISCUSSED WITH AI COMPANIES, INCLUDING ALIBABA, BYTEDANCE, AND ZHIPU AI, POSSIBLE RESTRICTIONS ON THE OVERSEAS TRANSFER OF KEY DATA USED TO TRAIN AI MODELS, AS WELL AS ON FOREIGN USERS DOWNLOADING MODEL WEIGHTS.
HOWEVER, CHINA PLANS TO CONTINUE ALLOWING OVERSEAS CUSTOMERS TO ACCESS THESE MODELS AND SERVICES.
Datacenters have always kept reciprocating engines on site as emergency standby, kicking in when the grid drops and running a handful of hours a year. That role is changing, as recips are increasingly being repurposed to prime power, energizing datacenters around the clock. We estimate that recip OEMs (e.g., Caterpillar, INNIO, Cummins) have been contracted to supply ~1 GW of BTM power this year, and 4+ GW in each of 2027 and 2028. Our SemiAnalysis Energy Model tracks these BTM contracts OEM-by-OEM, project-by-project. Yet, that is dwarfed by the opportunity to come. (1/3)🧵
PALO ALTO, California -- Hidden debt at U.S. tech giants swelled eightfold in roughly four years to an estimated $1.65 trillion as artificial intelligence investments ballooned, a Nikkei study shows, exceeding actual debt and making it tougher for investors to assess risk.
Nikkei examined recent financial statements and other materials from Google owner Alphabet, Microsoft, Amazon, Meta and Oracle. The four companies aside from Oracle are scheduled to announce their second quarter earnings from Wednesday, meaning the figures may increase further.
The five companies' hidden debt, which does not appear on balance sheets, totaled $1.65 trillion in the most recent quarter, exceeding the roughly $1.35 trillion in debt reflected on their balance sheets. The data includes some estimates.
‼️Chinese AI models have officially overtaken the US globally:
China's share of AI token usage by US firms on OpenRouter jumped to ~58%, surpassing the US for the first time since this data began.
This share has more TRIPLED over the last few months.
Among US firms specifically, Chinese models now account for ~38% to 40% of tokens used, with DeepSeek remaining the single most popular choice, well ahead of Z Ai model, Qwen, MiniMax, and Kimi.
Meanwhile, Chinese startup Moonshot released its new Kimi K3 model last week, a release investors say rivals top US systems and one that triggered a fresh selloff in AI and semiconductor stocks, echoing last year's "DeepSeek moment."
Furthermore, Chinese officials are reportedly weighing restrictions on foreign access to their most capable models, facing the same security dilemma already playing out in Washington.
The AI race is no longer just about who builds the most powerful models, but who can make them cheaper, faster, and more widely adopted.
We cut open the Kirin 9030 and put it under an electron microscope.
The smallest metal pitch measures 32.5 nanometers. That is tighter than Intel 18A, their brand new leading edge node.
A Chinese fab with no EUV is out-pitching Intel's EUV node by roughly 10 percent. What's going on here? 🤯
"This is the HiSilicon Kirin 9030, the chip inside Huawei's newest flagship phone. A few weeks ago, we cut it open, put it under an electron microscope, and measured the smallest wires inside the chip, the metal pitch. And what we saw was unexpected. The smallest metal pitch inside the Kirin 9030 measures only 32.5 nanometers. That’s smaller than the metal pitch in Panther Lake, which is based on Intel's brand-new 18A node.
A Chinese fab, cut off from the most advanced tools, without EUV, is packing its wires about ten percent tighter than Intel's leading edge EUV node. What’s going on here?"
🦔A Nikkei investigation found that Alphabet, Microsoft, Amazon, Meta, and Oracle have $1.65 trillion in debt that doesn't appear on their balance sheets, more than the $1.35 trillion they officially report. These are GPU contracts, data center leases, and joint ventures that don't count as debt under accounting rules until the facilities go live. Meta's hidden debt is $420 billion, triple its reported debt. Oracle's grew 30-fold in four years. All five declined to comment.
My Take
Nikkei examined the actual filings and put a number on something the BIS already flagged as "shadow borrowing" back in March. These companies owe more off their balance sheets than on them, and the accounting rules let them keep it that way until the data centers go live. That's legal, but it means investors looking at quarterly earnings this week are seeing less than half the picture.
Four of these five report earnings in the next two weeks. The reported debt will look manageable. The $1.65 trillion in footnotes won't make the headlines. But when those data centers start operating, the leases hit the books all at once. If AI demand comes in below projections, those facilities get marked down and the losses land on the investors and insurance policyholders who funded the construction through private credit and project bonds without realizing how much total exposure they were carrying.
Hedgie🤗
Kimi is striking while the iron’s hot
If Moonshot IPO’s anywhere near the $30B valuation they’re currently in the works on raising, it will likely shatter Anthropic and OpenAI’s plans to IPO at or above $1T (if it hasn’t already) regardless of their current ARR’s.
Valuations are of course forward looking. Current ARR is far less significant if there’s now another competing lab offering effectively the same caliber model for a fraction of the price.
US labs will either have to lower their pricing, or lose customers to Kimi and other competitive open-source models over time. Either way, ARR is bound to decline (with the exception being that the increase in demand for AI can offset the decline in pricing, but regardless that would slash margins)
The K3 release may not be bearish semis or infrastructure, but it certainly feels like an “emperor wears no clothes” moment for Anthropic and OpenAI.
Kimi K3 2.8T is so large that it will not fit on a single NVIDIA DGX B200, even at FP4. A GB300 NVL72, B300, or MI355X system is required, as each GPU has 288 GB of memory.
One optimization that could make Kimi K3 fit on B200 is to gang multiple nodes together and use a technique called WideEP. The issue is that B200 has only 400 Gbit/s of bandwidth between nodes, whereas NVL72 has 18× higher inter-node bandwidth.
Holy moly: Zhipu AI founder (GLM-5.2) Tang Jie says we are on our clear way to AGI and "AI will begin to learn what the "self" is and what self-awareness means"
In a purported internal letter, he argues that:
- autonomous agent systems are moving toward the fully automated “no-person company”: thousands of agents working continuously, collaborating, evaluating results and allocating resources.
- His more provocative claim: "AI training AI is already taking shape." (RSI) Models can increasingly write code, synthesize data and participate in training loops. Zhipu wants to push this further through self-play, synthetic-data factories and systems that can reconstruct their own code inside secure sandboxes, potentially generating new knowledge rather than simply recombining human output.
Long-horizon tasks → autonomous agent societies → fully automated “no-person companies” → AI training AI → self-evolution → self-awareness → emotion → consciousness → ASI.
Tang writes:
“AI will begin to learn what the ‘self’ is and what self-awareness means. Beyond that, it may begin to touch human emotion. Farther still lies consciousness itself.”
He believes memory, continual learning and self-evaluation - problems once thought to require an entirely new paradigm - are gradually being overcome.
Models are already beginning to write code, synthesize their own data and participate in training future models.
Zhipu now wants systems that can reconstruct their own code and generate knowledge through self-play.
Is that the beginning of recursive self-improvement?
Tang appears to believe so. His essay does not stop at more capable AI tools. It describes a direct progression from automated work to self-evolving intelligence, and eventually to machines that understand their own existence.
In short: today's LLMs will lead to ASI via AGI, context and memory will be solved, and AI will become self-aware.
I've rarely seen anyone write something so bullish. And if it weren't coming from the founder of GLM, I would dismiss it. But not only is he a true expert, but with GLM they've proven what they're capable of.
h/t @AndrewCurran_ He brought the essay to my attention.
Apple just sued OpenAI, and the wildest part is how they got caught: one candidate screenshotted confidential Apple files on his Apple work laptop hours before his OpenAI interview. Apple reads its own server logs. The recruiting pipeline generated its own evidence trail.
The complaint says OpenAI's hardware chief Tang Tan, a 24-year Apple veteran, directed candidates still employed at Apple to bring "actual parts" (batteries, logic boards) to interviews for show and tell sessions. One candidate was surprised, saying he didn't even know you could take those out of the office.
Apple also alleges Tan circulated an internal Apple offboarding document to coach new hires on dodging exit security checks, and that a departing engineer kept his Apple laptop, found a bug that still gave him access to Apple's cloud storage, and downloaded dozens of confidential hardware files after joining OpenAI.
Then the supplier: OpenAI allegedly got one of Apple's manufacturing partners to demonstrate a proprietary metal finishing technique by letting the partner believe Apple had approved it.
Over 400 former Apple employees now work at OpenAI. Apple says it flagged all of this to OpenAI in February and never got a response. Five months later, it filed.
The ask reveals the strategy. Apple wants an injunction barring OpenAI from using the secrets, the return of every file, and full discovery into io, right as OpenAI preps its first device launch and an IPO. If a judge grants it, OpenAI may have to prove the device was built clean, component by component, before it ships.
The device was supposed to run on the world's best hardware talent. Now its bill of materials is evidence.
The US economy is now dependent on AI spending:
AI investment now accounts for more than 25% of US GDP growth, the largest contribution on record.
This includes spending on software, IT equipment, R&D, and data centers.
In other words, for every $4 of US economic growth today, over $1 is coming from AI investment.
This comes as AI spending is up to a record ~8% of US GDP.
By comparison, spending on IT equipment, software, and R&D peaked at ~6.5% of GDP during the 2000 Dot-Com bubble.
US economic growth is now all about AI.
The AI infrastructure buildout is entering a new phase:
US tech companies are committing to spend a record $850 billion on data center leases over the next several years.
This marks a +$570 billion YoY increase, or +204%, and +$200 billion QoQ increase, or +31%.
Meta, $META, added the most in Q1 2026, committing +$79 billion in new leases, a +76% QoQ increase, bringing its total to ~$183 billion.
At the same time, Microsoft, $MSFT, added +$41 billion, a +26% QoQ increase, bringing its total to ~$197 billion.
Oracle leads with the largest total commitments at ~$250 billion, having already secured many of the key sites needed to fulfill its contract with OpenAI.
Tech companies are doubling down on AI.
Bloomberg: "The Silicon Data LLM Token Expenditure Index, which tracks what users pay for AI tokens, is down almost 20% from a high in May after nearly doubling since its inception in December. The gauge is the cleanest read anyone has on the $700 billion-plus capex boom that has done the sector’s heavy lifting. For stock investors, that could be flashing a warning that AI companies are losing pricing power with increasingly cost-sensitive customers, and that expectations for an eventual AI bonanza could prove misplaced.
'There are increasing reports that users of AI solutions, priced in tokens, are having to restrain unlimited use due to high costs,' said veteran investor Louis Navellier. 'The chatter that OpenAI is pushing back its IPO to next year is seen as a sign that, currently, profitability remains a problem.'"
As I warned back in my December report on "GenAI & Productivity" (https://t.co/kEx5Z4BbRz):
"GenAI appears to be transitioning tech giants from the most-profitable business models in history to business models more akin to industrials...Meanwhile, while there’s a lot of speculative fear about how a single LLM could rise to dominance and what that could mean for economic, societal, and political stability, we believe the bigger concern for investors today is how relative model parity could compromise pricing power. Tech giants have thrived on monopolies and duopolies for a decade or more. Now, they’re in an LLM arms race where it’s unclear when or even if ever leadership will be sustainable."
Learn more about Sage Road Research here: https://t.co/Wgwz2xmY1y. Interested in subscribing? Message me.
Bloomberg link: https://t.co/dmiMU860Oh