KV cache and model weights are the two big memory consumers in modern LLMs.
Weights = fixed footprint, shared across every user.
KV cache = grows with every conversation's length AND every concurrent user.
At frontier scale, KV cache is the real bottleneck, not weights.
Why: weights load once and amortize across millions of users. KV cache is marginal. it multiplies with concurrency × context length. Long chats + lots of users = KV cache dominates GPU memory, not the model itself.
Two techniques attacking this right now:
Google's TurboQuant (ICLR 2026): compresses KV cache to 3-bit, near-lossless, ~6x smaller
Kimi Linear: swaps most full-attention layers for linear attention — replacing an O(n) cache with a fixed O(1) state. Interleaved 3:1 with full attention to keep retrieval quality, cuts KV cache ~75%
Stack both and you can push KV cache memory down to a small fraction of the FP16 baseline at long context , turning what used to be a hard memory wall into something manageable.
A lot of innovation happening at very fast pace.
MERITZ SECURITIES: ACCORDING TO CHANNEL CHECKS, MIDDLE EASTERN SOVEREIGN AI INVESTORS, INCLUDING THOSE IN SAUDI ARABIA, HAVE RECENTLY BEGUN DISCUSSING MID- TO LONG-TERM MEMORY PROCUREMENT PLANS WITH KOREAN MEMORY MANUFACTURERS.
- Amid this demand growth, upward pressure has begun to emerge in the server DRAM spot market.
- Prices are rising particularly sharply for high-end products, such as those with 6,400Mbps bus speeds.
- This suggests that intensifying investment competition among CSPs and frontier-model developers is increasingly focused on products designed to maximize performance, amid worsening supply shortages.
- Spot prices for 64GB DDR5 server DRAM products have risen steeply since mid-July.
- Recent prices have climbed to $3,100–$3,400, approximately 146% above the end-June contract price of around $1,380.
- ACCORDINGLY, SERVER DRAM CONTRACT PRICES IN 3Q26 COULD RISE BY MORE THAN THE MARKET’S CURRENT EXPECTATION OF APPROXIMATELY 15% QOQ DURING THE QUARTER.
- In particular, suppliers that adopted more customer-friendly and flexible pricing in 2Q26 are likely to see especially sharp price increases in 3Q26 and 4Q26.
@sheeprep@HiCagr How is KDA bearish NAND? It’s the polar opposite. This will be DeepSeek moment where the momentum sell off has started through memory.
Once this “realization” permeates it’ll get bought back up because everyone comes to the opinion “oh it’s great for sandisk eSSDs”