This Warsh contradiction has been nagging at me. At Sintra at the start of the month, he took comfort at the recent decline in bond yields, implying bond markets understood low inflation was on the way. Today, he took comfort at *higher* bond yields, saying they will deliver low inflation. How can this be: that lower bond yields are reasons to feel good about inflation, but higher bond yields are not a reason to feel bad about it? Without him articulating a monetary and economic framework, these statements make it feel like he's winging it.
A guide to Warsh's rhetorical techniques. (since we'll living with it for the next 5 years)
Here's 8: :
1. Replace the question with a broader one (otherwise known as "bridging), i.e. when someone ask X, you answer Y. Q: Was today's decision close? Answer: The vote was 9-3. The broader discussion showed a lot of agreement. We had commonality on the questions....
2. Reject the reporter's premise. I.e. Q: What's the argument for a pause? Answer: I wouldn't characterize what we did as anything like a pause.
3. When asked about conclusions, answer with process. Q: What haven't you hiked? A: We spent an inordinate amount of time looking at monetary policy strategy...thinking hard...looking at data...
4. Admit uncertainty, but only about the economy, not objectives. i.e. he never expresses uncertainty about the 2% inflation target or the Fed's mission, but he freely admits uncertainty about anything about the economy.
5. Create memorable slogans to create rhetorical anchors. i.e. "Watchful thinking, not watchful waiting." "Family fight." "We're in the performance business." "No magic wand."
6. Controlled self-deprecation to diffuse tension. i.e. "Believe it or not, this press conference is not all I've done today."
7. Use questions to answer questions. He presented the four questions that the FOMC discussed, but never provide answers.
8. Divert attention to the markets but does not provide market analysis. Q from reporter: what are you hearing from the markets? A: Markets are the best source of information.
The chart is visually powerful, but economically incomplete because every arrow is treated as though it represents the same type and degree of risk. It does not. An equity investment, a cloud-purchase commitment, a capacity reservation, a supplier warrant, a revenue-sharing arrangement and a debt guarantee may all create interconnected exposure, but they have very different consequences for cash flow, solvency and revenue quality. The correct conclusion is therefore not that the entire AI ecosystem is fraudulent or that every dollar of reported demand is circular. The more defensible conclusion is that AI infrastructure is increasingly being financed through a reflexive system in which suppliers fund customers, customers commit to buying from suppliers, infrastructure providers borrow against those commitments and rising private-market valuations make it easier to raise the next round of capital.
Circular financing is not automatically problematic. In a market where demand is expanding rapidly, advanced chips are scarce and the infrastructure required to operate them costs tens of billions of dollars, strategic financing can solve a genuine coordination problem. The company developing the model may have the technological capability and future revenue opportunity but lack the capital required to build infrastructure today, while the chipmaker or cloud provider has the balance sheet, equipment and strategic incentive to accelerate that buildout. Financing and purchase commitments can therefore bring forward productive capacity that would otherwise arrive too slowly. OpenAI itself describes compute, distribution and capital as the three requirements for scaling AI, and its February 2026 financing included $30 billion from $NVDA, $50 billion from $AMZN and $30 billion from SoftBank alongside new infrastructure partnerships.
The risk begins when capital supplied by the ecosystem becomes difficult to distinguish economically from demand generated by the ecosystem. A supplier investing in a customer does not automatically invalidate the revenue that follows, provided real equipment is delivered, the customer receives economic value and the obligation is collectible. However, it does weaken the evidentiary value of the order. An order funded from internally generated cash flow tells investors that end customers have already monetized sufficient demand to pay for the infrastructure. An order funded by the supplier’s equity investment, debt guarantee or credit support tells investors that the market expects sufficient demand to emerge later. Both can eventually become profitable, but they represent different levels of commercial validation.
That distinction is central to understanding the current AI cycle. The most important question is not whether Nvidia invests in a company that buys Nvidia chips. The more important question is where the money ultimately originates. If enterprises and consumers are paying enough for AI products to finance model development, cloud consumption and hardware depreciation, then the capital flows are merely accelerating a legitimate adoption cycle. If the AI laboratories are paying their infrastructure bills primarily with money raised from the same companies supplying that infrastructure, then the industry may be temporarily financing its own revenue while waiting for external monetization to catch up.
The ecosystem is also not equally fragile at every layer. $MSFT, $GOOG, Amazon and $META can fund enormous AI investments from established businesses spanning enterprise software, advertising, commerce and cloud computing. Even if returns on individual AI projects disappoint, these companies possess diversified operating cash flows and investment-grade balance sheets. OpenAI, Anthropic and other frontier laboratories sit in a more vulnerable position because their compute obligations can grow faster than their current revenue and because they rely much more heavily on external capital. The greatest financial fragility may sit one layer below them in highly leveraged infrastructure providers that borrow heavily to build data centers against contracts concentrated among only a handful of AI customers.
$CRWV illustrates both the strategic logic and the financial risk. Nvidia invested $2 billion in CoreWeave in January 2026, while the companies agreed to collaborate on more than five gigawatts of AI-factory capacity by 2030. CoreWeave has disclosed that all GPUs used in its infrastructure are Nvidia GPUs as a result of its customer obligations, while Nvidia also has contractual access to residual unsold CoreWeave capacity under a multibillion-dollar agreement. At the same time, CoreWeave historically generated a very large share of its revenue from Microsoft, with Microsoft accounting for approximately 67% of revenue in the third quarter of 2025, and the company has funded growth through substantial amounts of secured and unsecured debt, including bonds issued at interest rates around 9%.
None of that proves the demand is artificial. CoreWeave provides real computing capacity, Microsoft and other customers use that capacity and Nvidia delivers physical hardware with genuine economic utility. However, the structure concentrates several risks inside the same network. If model demand weakens, the AI laboratory may reduce cloud consumption; the infrastructure provider then faces lower utilization while retaining debt, lease and power obligations; Nvidia loses future chip demand and may also suffer losses on its equity investment or contractual exposure. The same shock therefore travels through revenue, asset values and credit simultaneously.
The largest risk to Nvidia is not necessarily the direct loss on its investments. Nvidia reported fiscal 2026 revenue of $215.9 billion, while its April 2026 filing disclosed approximately $27 billion of investment commitments and $18.6 billion invested in private companies and infrastructure funds. Those sums are meaningful but manageable relative to Nvidia’s earnings power and market capitalization. The more serious risk would arise if strategic financing has brought forward several years of GPU purchases that end-user AI revenue cannot ultimately support. In that scenario, Nvidia would not merely impair an investment; it could experience a decline in future orders precisely when customers are attempting to repair their balance sheets.
This is why the circularity debate should focus less on accounting and more on economic substance. Equity investments are recorded separately from product revenue, and genuine hardware deliveries do not become fraudulent merely because the supplier owns part of the customer. There is no public evidence that the current AI partnerships are equivalent to sham round-trip transactions. Nevertheless, economic circularity can exist without accounting fraud. Revenue may be properly recognized under accounting standards while still depending on capital provided by related ecosystem participants rather than independently generated customer cash flow.
The late-1990s telecommunications comparison is useful, but only when applied precisely. The internet was real, bandwidth demand did grow and the infrastructure eventually became essential. Investors nevertheless lost enormous amounts of money because the industry built too much capacity too quickly, financed weak customers and assumed that exponential traffic growth would automatically produce attractive returns on capital. When bandwidth prices collapsed, heavily leveraged carriers cut equipment spending, excess fiber remained underutilized and losses propagated back toward equipment suppliers and creditors.
Some telecom transactions went considerably further and crossed into improper accounting. The SEC found that Qwest entered capacity swaps in which it bought capacity it did not need while counterparties bought capacity from Qwest, sometimes supported by concealed agreements that undermined the claimed revenue recognition. That is materially different from Nvidia selling functional GPUs to an operating data center or Microsoft selling genuine cloud capacity to an AI developer. The historical lesson is therefore not that every reciprocal commercial relationship is fraudulent. It is that genuine technological demand can coexist with excessive financing, distorted incentives and, in the weakest cases, attempts to manufacture revenue.
There is another important difference between fiber and AI infrastructure. Fiber can remain physically usable for decades, but individual semiconductor generations depreciate technologically much faster. A data center built around today’s accelerators may face newer chips offering substantially better performance per watt and lower cost per token before the original equipment has earned its required return. This creates a risk that the economic life of the hardware is shorter than the financing period attached to it. However, obsolescence is not binary. Older GPUs can migrate from frontier training toward inference, fine-tuning, research and lower-cost workloads, while power connections, cooling systems, land and data-center shells can retain value even after the computing equipment is replaced.
The question is therefore whether utilization and pricing decline faster than the owner can depreciate the equipment and refinance the liabilities. A facility operating at high utilization can generate attractive cash flow even if the hardware becomes second-generation technology. A facility operating below capacity while rental prices fall and interest expense remains fixed can destroy equity very quickly. The largest vulnerability is not technological obsolescence by itself; it is technological obsolescence combined with leverage, customer concentration and long-dated take-or-pay obligations.
Microsoft and OpenAI also show that interconnectedness can evolve rather than remain permanently circular. Microsoft remains a major OpenAI shareholder and reported an approximately 27% ownership interest on an as-converted basis, while Azure continues to play a central infrastructure role. However, the partnership was amended in April 2026 to make Microsoft’s intellectual-property license non-exclusive, allow OpenAI to serve products across other cloud providers and remove Microsoft’s payment of a revenue share to OpenAI, while OpenAI’s revenue-sharing payments to Microsoft continue through 2030 subject to a cap. This diversification reduces OpenAI’s dependence on one supplier, but it also spreads the ecosystem’s exposure across more cloud providers and chipmakers rather than eliminating it.
Supporters are therefore correct that many of these deals constitute a rational form of industrial coordination. Building frontier AI infrastructure requires aligning semiconductor production, networking, data-center construction, electricity supply, model development and customer distribution several years before demand is fully visible. Traditional spot-market purchasing cannot efficiently coordinate a buildout of that scale. Long-term commitments, strategic equity and co-investment can lower financing costs, secure scarce capacity and distribute risk among parties that benefit from the ecosystem’s growth.
Critics are equally correct that this structure can weaken price discovery. A supplier that owns equity in its customer may accept commercial terms that an independent supplier would reject. A cloud company with a large investment in an AI laboratory may continue providing capacity because protecting the equity value becomes part of the decision. An infrastructure provider may build against commitments from counterparties whose ability to pay depends on future fundraising rather than current cash generation. The danger is not simply that participants make irrational decisions. It is that individually rational decisions, each designed to protect an existing investment or strategic relationship, collectively sustain uneconomic capacity for longer than an independent market would.
The most revealing metric is therefore not announced investment or contracted backlog in isolation. Investors need to determine how much of the backlog is supported by investment-grade counterparties, how much is cancellable, whether the contracts contain take-or-pay protection, whether the customer has generated the cash independently, whether equipment can be redeployed and whether the infrastructure owner’s debt maturity is shorter than the expected monetization period. A $20 billion commitment from a cash-rich hyperscaler is fundamentally different from a $20 billion commitment by a loss-making laboratory whose ability to pay depends on raising its next $50 billion funding round.
Cash conversion will become increasingly important. Reuters estimates that investment by the major hyperscalers is rising considerably faster than operating cash flow, with approximately $534 billion of incremental capital expenditure expected by 2027 against roughly $340 billion of incremental operating cash flow. That does not by itself imply overinvestment because infrastructure spending is front-loaded while revenue arrives later, but it raises the burden of proof. The longer capital expenditure grows faster than operating cash generation, the more the investment case depends on future AI revenue rather than demonstrated current economics.
The bullish case remains credible because end demand is already broadening. AI is moving beyond model training into coding, advertising, search, enterprise automation, scientific research and persistent inference, while newer reasoning and agentic workloads can consume substantially more compute than conventional chatbot interactions. If falling token costs expand usage faster than efficiency reduces the compute required per task, the industry can grow into much of the infrastructure currently being built. Under that outcome, circular financing will be remembered as the bridge that allowed the supply chain to scale ahead of demand.
The bearish case is not that nobody uses AI. The bearish case is that AI becomes ubiquitous while the financial returns on today’s infrastructure remain poor. The internet transformed the world, but many companies that financed its initial infrastructure still failed. Technological importance does not guarantee that every participant earns its cost of capital. AI models may become extraordinarily valuable while competition pushes model prices down, open-source systems compress margins, hardware improves faster than assets depreciate and the largest share of the economic profit accrues to a narrower group of platforms than today’s investment boom assumes.
My view is that AI is not a repetition of the telecom bubble in the simplistic sense that demand is imaginary or the technology lacks utility. The demand is real, the productivity potential is substantial and the leading hyperscalers possess far stronger balance sheets than the speculative telecommunications carriers of the late 1990s. However, parts of the financing structure are increasingly reminiscent of vendor-financed infrastructure booms, particularly where suppliers provide equity, guarantees or capacity commitments to entities that then become major purchasers of the suppliers’ products.
The circularity itself will not necessarily end the AI boom. It will determine where the losses concentrate if the boom slows.
I remain structurally bullish on AI infrastructure, but the network of reciprocal investments argues for greater selectivity rather than blanket enthusiasm. The strongest positions remain businesses with technological bottlenecks, pricing power, diversified customers and substantial internally generated cash flow. The most vulnerable positions are highly leveraged infrastructure providers with concentrated customers, fixed power and lease commitments, rapidly depreciating hardware and counterparties whose purchasing power depends on continuing access to capital markets.
The warning signs would be declining GPU utilization, falling rental prices that cannot be offset by lower hardware costs, repeated contract renegotiations, customers delaying deployments, suppliers providing progressively larger guarantees to preserve orders and AI laboratories raising capital primarily to meet existing infrastructure obligations rather than funding new growth. The most constructive signals would be the opposite: enterprise AI revenue expanding rapidly, inference utilization broadening, customer concentration falling, free cash flow improving and infrastructure commitments increasingly being paid from operating revenue rather than supplier-backed financing.
The sharpest conclusion is that the AI ecosystem is not currently a fraudulent circle, but it is becoming a financially reflexive one. Reflexivity magnifies the upside while capital is abundant, asset prices are rising and demand exceeds supply. It also magnifies the downside when one of those conditions reverses.
A real technological revolution can still produce an investment bubble around parts of its infrastructure. The internet proved both propositions simultaneously. AI may do the same.
If the Fed hikes this week, is that bearish bonds because the Fed is more worried about inflation than we thought, or bullish because they are getting ahead of the curve? If only there was some way they could tell us, you know, through some form of foward-looking guidance. Amazing that no CB has ever come up with a solution to this problem. Oh well...😉
CXMT's $CXMT IPO is another reminder that the memory story is no longer simply about supply. It is increasingly becoming a story about demand.
Many investors will look at the emergence of China's fourth major #DRAM manufacturer and immediately conclude that this is bearish for #memory pricing. More capacity means more competition, lower prices and weaker margins. That has certainly been true throughout much of DRAM's history.
We think #AI fundamentally changes that equation. The biggest takeaway from CXMT's remarkable debut is not that another competitor has entered the market. It is that memory has become one of the most strategic assets in the AI value chain. Investors are assigning a valuation approaching half a trillion dollars to a company that, just a few years ago, barely registered in the global semiconductor industry. Markets rarely assign those valuations to businesses they believe are entering structural oversupply.
More importantly, the demand backdrop today looks nothing like previous memory cycles. Traditional DRAM demand was largely driven by PCs, smartphones and servers, all relatively mature end markets with modest memory content growth. AI changes that entirely. Every frontier model requires enormous amounts of memory during training, while inference is becoming increasingly memory intensive as context windows expand, reasoning models become larger and agentic AI performs more complex tasks.
We continue to believe the market is underestimating how large the AI memory opportunity ultimately becomes. Many investors still think in terms of ChatGPT replacing Google searches or AI assistants helping employees write emails. That is only the first chapter. The much bigger opportunity emerges when AI evolves from answering questions to performing work.
An agent that writes code, books travel, negotiates contracts, manages supply chains, analyzes medical images or controls robots consumes significantly more compute and memory than today's chatbot interactions. If millions or eventually billions of AI agents begin operating continuously rather than responding only when prompted, inference demand could increase by an order of magnitude.
That is where the memory story becomes particularly compelling. Every reduction in inference cost expands the addressable market. Cheaper intelligence encourages more applications, more users and longer usage. History consistently shows that efficiency improvements do not reduce infrastructure demand. They expand it. This is classic Jevons Paradox. Cloud computing made servers cheaper to access. Demand for servers exploded. Internet bandwidth became cheaper. Global internet traffic exploded. Storage costs collapsed. The world stored exponentially more data. We believe AI will follow the same pattern.
Even if model architectures become more efficient through techniques such as Mixture-of-Experts, quantization or KV cache optimization, those improvements lower the cost of intelligence rather than reducing demand for intelligence. As AI becomes cheaper, entirely new applications become economically viable, expanding total compute demand and, by extension, memory demand. This is why we increasingly think the memory pie is substantially larger than consensus expects.
CXMT's rise also reinforces another important point. The memory market is no longer constrained solely by manufacturing capacity. It is increasingly constrained by geopolitical fragmentation. China wants domestic supply. The United States wants trusted supply chains. Apple wants additional suppliers. Governments want technological sovereignty.
These objectives do not necessarily compete with one another. They often require parallel investment across multiple regions, leading to more capacity being built than would exist under a fully globalized market. In other words, geopolitical fragmentation may actually increase total capital spending across the memory industry.
From an investment perspective, we continue to distinguish between conventional DRAM and AI memory. Commodity DRAM will likely remain cyclical over time as Chinese capacity expands. HBM is a different business. HBM requires advanced packaging, extremely high yields, sophisticated thermal management and close integration with GPU vendors. Those capabilities remain concentrated among a handful of companies, particularly SK Hynix, with Micron rapidly catching up. While CXMT is making impressive progress, its HBM3 efforts are still several years behind the industry's technology leaders.
Ultimately, we think investors should avoid viewing this as a zero-sum story where China's gain automatically becomes everyone else's loss.
The more important observation is that every major technology company, every hyperscaler and now every major government is investing aggressively in AI infrastructure. That tells us the bottleneck is no longer whether AI will be adopted. The bottleneck is building enough infrastructure to support that adoption.
Our long-term thesis therefore remains unchanged. We believe AI is creating a structural rather than cyclical expansion in memory demand. If AI adoption continues accelerating and agentic AI becomes mainstream over the next decade, the total addressable market for memory could prove substantially larger than most investors currently model. In that environment, multiple winners can coexist. The pie itself is becoming much larger.
Where's FCNR(B) flow gone? @kantisoumya of @TheOfficialSBI says, they are in transit.
So far (8 June-17 Jul), FCA up by $7.6 bn, while RBI deposit mobilization was $17.4 bn for FCNR(B) flows.
While RBI provides weekly change in FCA every Friday, every bank has a designated time during the week to swap these FCNR(B) deposits with RBI and get equivalent rupee resources. It's possible that FCNR(B) deposit numbers are being captured by RBI and becomes a part of FCA with a lag --with assumption that a large part of foreign exchange reserves is being recouped & adding to RBI's foreign exchange coffers.
Owing to this, we expect for the next 15 days period (17 Jul-31 Jul) FCA inflows could touch $10-12 billion based on current trends.
It may also be noted that once the swap is done, there will be a corresponding liquidity injection (permanent and temporary) Thus, bank time deposits/ permanent monetary injection have increased by close to Rs 90,000 crores. Core liquidity / Temporary liquidity injection has already at ~Rs 5 lakh crore as of June 30 clearly indicating the impact of FCNR (B) inflows on liquidity.
Our recent research is now available in one- to two-minute videos.
From AI and manufacturing to trade, competition, and America’s competitive edge, hear directly from our authors on the ideas shaping the global economy.
Subscribe to Forward Thinking:
https://t.co/EueWGs2s5I
Nomura's report reinforces what we believe is one of the most underappreciated structural shifts taking place across the semiconductor industry today, namely that testing is no longer a low-value manufacturing step performed at the end of production, but is rapidly evolving into one of the most critical value-added processes within the entire AI hardware supply chain, because as chips become exponentially more complex through chiplets, stacked HBM, advanced packaging, silicon photonics and eventually co-packaged optics (CPO), the economic cost of failure rises disproportionately, making every additional dollar spent on testing significantly more valuable than it was during previous semiconductor cycles.
The market has understandably spent the past two years focusing almost exclusively on GPU designers, HBM suppliers and advanced packaging companies, yet what this report demonstrates is that testing is quietly becoming the next bottleneck, because increasingly sophisticated AI accelerators cannot simply be manufactured, they must be validated repeatedly throughout the production process to ensure every component performs flawlessly before being assembled into AI systems that may ultimately be worth several million dollars each, effectively transforming testing from a manufacturing support function into an essential yield protection mechanism.
Historically, testing was largely viewed as a necessary manufacturing expense whose primary objective was to filter out defective chips before shipment, but AI has fundamentally altered that equation because testing today is increasingly about protecting economic value rather than merely measuring quality, and when a single package contains multiple GPU chiplets, twelve stacks of HBM, advanced substrates, hybrid bonding interfaces, silicon photonic engines and increasingly expensive packaging materials, discovering a defect late in the production process can destroy vastly more value than in previous semiconductor generations.
That is precisely why Nomura estimates testing content continues to increase materially with every GPU generation, using Hopper as the baseline, where final testing time increases approximately fourfold for Blackwell and roughly sevenfold for Rubin, while system-level testing rises approximately 1.5 times for Blackwell and 2.5 times for Rubin, with burn-in testing roughly doubling, resulting in testing content increasing from approximately 1.9% of total GPU cost for Hopper to 2.5% for Blackwell and approximately 3.3% for Rubin, a progression that may appear modest when expressed as percentages but becomes extraordinarily meaningful when applied to AI systems whose selling prices continue rising dramatically.
Perhaps the most important observation in the report is not simply that testing content is increasing, but that testing itself is migrating earlier throughout the manufacturing process, effectively shifting from a single inspection performed after fabrication into a continuous validation framework that begins at wafer probing, continues through known-good-die verification, hybrid bonding validation, package testing, burn-in qualification and ultimately system-level testing before deployment inside hyperscale AI clusters, meaning the industry is increasingly adopting multiple quality gates rather than relying on one final inspection at the end of production.
This shift has enormous implications for the supply chain because every additional testing insertion creates incremental demand for specialized equipment, probe cards, sockets, handlers, MEMS probes, thermal management systems and high-speed interfaces, thereby expanding the opportunity set well beyond traditional outsourced semiconductor assembly and test companies, which explains why Nomura has broadened its coverage to include interface suppliers and test hardware manufacturers rather than limiting its investment thesis solely to OSAT providers.
Another theme that deserves significantly more attention is the interaction between advanced packaging and testing, because while investors have understandably focused on CoWoS capacity as one of the industry's largest bottlenecks, packaging capacity alone cannot solve the industry's challenges if testing capacity fails to expand at a similar pace, since every additional layer of complexity introduced through chiplets, hybrid bonding, HBM stacking, heterogeneous integration and silicon photonics simultaneously increases the probability that expensive failures will occur after substantial value has already been added to the product, making testing increasingly indispensable as AI hardware becomes more sophisticated.
The discussion surrounding co-packaged optics is equally compelling because most investors naturally associate CPO with optical component suppliers, whereas Nomura correctly argues that the real opportunity extends much further into the testing ecosystem, given that every optical engine must communicate flawlessly with adjacent ASICs under extremely demanding thermal, electrical and optical conditions while maintaining signal integrity across increasingly complex architectures, thereby introducing entirely new categories of testing that simply did not exist in previous semiconductor generations and creating an additional secular growth driver for testing vendors.
We also agree with Nomura's conclusion that the AI infrastructure cycle remains considerably earlier than many investors assume, because every successive GPU generation is becoming disproportionately more difficult to validate than its predecessor, allowing testing content to grow materially faster than semiconductor unit volumes themselves, which means the industry's next major beneficiaries may not necessarily be the companies designing the chips, but increasingly the companies ensuring those chips actually function reliably inside increasingly expensive AI systems.
Our preferred way to position for this theme is to own the entire testing value chain rather than focusing solely on OSATs, because different parts of the ecosystem benefit from different stages of the testing process. ASE (3711 TT) remains our highest-conviction OSAT exposure given its scale, broad customer base and dominant position across advanced packaging and testing. Hon Precision (7769 TT) stands out as one of the most attractive pure-play beneficiaries of final testing and system-level testing, areas where AI complexity is expanding the fastest. WinWay (6515 TT) offers differentiated exposure through sockets and probe cards, which should experience rising content per AI accelerator as electrical and thermal requirements become increasingly demanding. MPI (6223 TT) is well positioned through wafer probing, benefiting directly from the industry's shift toward earlier testing insertions and known-good-die validation. KYEC (2449 TT) remains an attractive second OSAT exposure with meaningful leverage to AI testing demand, while Chroma (2360 TT) provides exposure to automated testing equipment, allowing investors to participate in the hardware upgrade cycle required to support increasingly sophisticated AI devices.
Ultimately, we believe semiconductor testing has quietly transitioned from a manufacturing support function into one of the industry's most valuable strategic chokepoints, and just as HBM suppliers, advanced packaging companies and foundries have enjoyed structurally stronger pricing power because they occupy indispensable positions within the AI value chain, testing vendors increasingly appear poised to achieve similar economics as AI hardware becomes more heterogeneous, more thermally demanding, more optically integrated and substantially more expensive to manufacture, making testing one of the highest-conviction secular investment opportunities across the broader semiconductor ecosystem.
Moonshot AI just dropped a technical report explaining why Kimi K3 became the strongest open coding model
The shift: bigger models aren’t the future
Sparse Routing → Delta Attention → Attention Residuals → 1M Context → Efficient Scaling
Kimi K3 is built around five ideas:
• Sparse Routing: 896 experts exist, but only 16 are activated for every token. Most of the model stays idle.
• Delta Attention: instead of recomputing the entire context, attention focuses only on what’s changed. Lower compute, longer context.
• Attention Residuals: preserve useful information across layers instead of rebuilding representations from scratch.
• 1M Context: entire repositories, books, and research papers fit into a single session.
• Efficient Scaling: frontier coding performance without frontier inference costs.
Dense models execute every parameter
Kimi K3 executes only the experts that matter
That single architectural decision changes the economics of frontier AI
This technical report changed how I think about scaling open models
Read it first, then explore the article below
So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we've seen before, and did it at record speed.
Also massively grateful to @Zai_org: they shared GLM5.2 as open weights (for free!) with the world and it became a key part of our defense.
This is day one for cybersecurity in the age of agents & we're all learning that secrecy is not the answer & that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!
Companies invest where they expect to compete successfully and can make a decent return.
Investment has diverged globally, stalling in Europe, shifting in the US, and increasing markedly in China.
See where investment happens and why: https://t.co/kW24qZkHC3
Super interesting!
"Dollarisation and monetary control: what lessons for the rise of stablecoins?" by Boris Hofmann, Aaron Mehrotra, and Jan Paulick.
"The emergence of stablecoins has created a new channel to access US dollar liquidity in emerging market and developing economies (EMDEs), similar to the historical role of foreign currency deposits, or "deposit dollarisation". This has raised concerns about the possible implications for monetary control in EMDEs. Drawing on data on foreign currency deposits and dollar-pegged stablecoin inflows for more than 130 economies, we compare the dynamics and drivers of "stablecoin dollarisation" with those of conventional deposit dollarisation. We document that historical deposit dollarisation and recent stablecoin flows are both associated with similar macro-financial drivers, including the strength of exchange rate pass-through and sovereign or banking crises. We further document significant persistence in both deposit and stablecoin dollarisation, suggesting that dollarisation is hard to reverse once established. Unlike deposit dollarisation, stablecoin flows seem to be largely unaffected by either broad or specific capital flow restrictions. This likely occurs because stablecoins are partly circulating outside the regulatory perimeter. The historical record also suggests that moderate deposit dollarisation has been associated with somewhat higher inflation risks, although there is little evidence of significant impacts on monetary policy transmission."
https://t.co/GZkO9pNbmh
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
Andrej Karpathy just challenged one of the biggest assumptions in AI.
Most people think better agents come from bigger models.
He argues the opposite.
In 16 minutes, he explains why small models, the right tools, and closed feedback loops are where real capability comes from.
Watch it.
Bookmark it.
Easily the most impressive thing I saw at WAIC, and honestly the main reason I went to the conference in the first place. I see robots and AI companies all the time. I don’t get to see SuperPoDs very often.
This is Huawei’s new Atlas 950 SuperPoD behind me & @pstAsiatech. Also yes, it’s spelled PoD, not “pod.” It refers to point of delivery and comes from networking terminology.)
1,024 Ascend 950 chips in one turnkey AI system. Huawei says the architecture scales to 8,192 chips as a single logical computer, and then you can build clusters of these to half a million chips. The scale is kind of nuts.
Also noticed basically everyone is selling some version of this now. Huawei has SuperPoD. MetaX and Biren have SuperNodes. NVIDIA has DGX SuperPOD and NVL72. ZTE has its own integrated AI computing platform.
Nobody’s really selling chips anymore. They’re selling the whole AI computer. Chips, networking, interconnect, software, management … everything is integrated
.@TheEconomist "Ashoka" columnist: "As other Asian countries saw energy prices surge, mandated working from home and, in the case of the Philippines, declared a national emergency, most Indian households barely noticed any disruption. How did India do it? https://t.co/5YrydeGzT8
Quick reactions to the Houthi blockade on Saudi Arabia announced this morning, just before U.S. markets opened.
TL;DR: What matters isn’t just lost oil exports from Yanbu, but also lost Saudi and GCC grain IMPORTS through the Red Sea.
Basically, Iran and the Houthis are threatening the Saudi economy with exactly the sort of broad-based economic warfare the U.S. is subjecting Iran to through its blockade.
There are 5 major Saudi ports in the Red Sea. From north to south they are Duba, Yanbu, King Abdullah, Jeddah and Jazan. Southern ports are most exposed to Houthi attack.
Yanbu sends 4.5 mb/d of oil into global markets. If Houthi attacks shut down operations for a meaningful span, that’s obviously a big blow to markets.
But since the Iran War shut down Hormuz, Saudi’s Red Sea coast has become ever more crucial for dry goods imports like wheat, rice, consumer goods, etc., flowing not just to Saudi markets but other Gulf countries, too.
Saudi Arabia has massive wheat stockpiles, enough for 4-6 months of normal consumption. But they don’t have large reserves of rice or feed grains, meaning that other pantry staples and dairy/meat production could be more affected.
It’s highly unlikely that anyone in Saudi or the Gulf countries will starve if food imports are curtailed, even for months.
But it will make like unpleasant for those living there. Scarcity will push local inflation — prices of food and consumer goods will go up, putting pressure on the government to change its policies.
Does that sound familiar? Again it’s the basic strategy the U.S. is trying to pressure Iran, without much luck so far.
One other thought: It’s noteworthy that the blockade is targeting Saudi Arabia specifically, not international shipping through the Bab-el-Mandeb strait.
The targeted approach suggests to me that Iran and the Houthis are trying to keep a handle on escalation, and to show restraint but also hold capacity in reserve to target all shipping should the conflict escalate further.
Weekly
China's combination of recovering output, weak expenditure, and mixed monetary isn't great, but also doesn't suggest a big change in macro policy. I think it unlikely that Japan's bid for repatriation works without more BOJ normalisation. Korea and Taiwan cycles remain strong, implying rate hikes.
https://t.co/stt8frRRZi