We just released a massive update on our gpu performance engineering resource list
AI Performance Engineering v2
this is the most comprehensive resource list for learning gpu and ai perf engineering
link in thread 🧵
this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization.
it's an opinionated list on what we think is important to deeply understand how to optimize inference systems.
also added:
- flashattention-4, blackwell tensor memory, and low-precision tensor cores
- continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding
- moe serving, collectives, topology, and prefill/decode disaggregation
- blackwell ultra, mi350/cdna 4, ironwood, and trainium3
- kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5
Compute Strategy for VCs: AI-focused VCs raising funds now should 1.5-2x the amount they are targeting and spend the additional on reserved compute facilities that they allocate to portcos. Forward-thinking VCs are already doing this for late-stage funds and this strategy is starting to be adopted by growth and early-stage funds.
Founders: If your lead VC cannot find you compute, that's not a lead, that's a co-investor.
LPs: If a VC says they are investing in AI startups, ask them their compute strategy.
1. AI portcos will spend most of their money on compute anyway, so it is more efficient for VCs to directly acquire compute and benefit from economies of scale (pricing power, hard-to-acquire tacit knowledge) instead of each portco wasting time on it.
2. Compute will be very hard to find next year, so having compute becomes a serious capital differentiator i.e. access advantage, at the same time that most venture capital is becoming commoditized due to e.g. SPVs, crossovers.
3. Compute will likely become more expensive, so reserving compute now and offering "$50M worth of compute" in 6, 12 months may actually cost VCs much less than that.
4. VCs are positioned to bridge the creditworthiness gap between datacenters (i.e. their financiers) and startups. Because startups are new and unknown, they have to pay worse economic terms for GPUs and are often pushed to the back of the line for strategic reasons. The way that most startups get around this is to raise more money and lean on the prestige of their VCs. Consequently, VCs can and should make compute cheaper for their startups by loaning their prestige and name to compute procurement, directly.
Example fund math and how to allocate a cluster:
A serious early-stage fund should be raising $100-$200M today (If you want to do Seed and A in AI with less than that you are kidding yourself). Out of a $200M early-stage fund, perhaps 40% should be reserved for initial capital, 30% for follow-on, and 30% for compute (I'm eliding fees and overhead). That gets you $60M for compute. At a fictive price of $4.8 / GPU-hour for B300s, that's just about enough for a 64 node B300 cluster (512 GPUs) for 3 years.
With a 64 node / 512 GPU cluster you should plan to allocate roughly like this (illustrative):
* 1x 32 node block dedicated to large training jobs, held for 3-6month chunks
* 1x 16 node block as a bridge cluster for R&D / training, held for 3 month chunks
* 3x 4 node blocks as bridge clusters for inference, held for 1 month chunks
* 4x 1 node blocks as flex blocks for compute grants to potential founders, nonprofit support, "spare change under the couch cushion" type capacity for portcos, held for 6-8 week chunks.
Depending on your portfolio needs, you will want to adjust this schedule, but you want to hold a very large subset of the GPUs aside for large training jobs and as bridge capacity, a small amount as flex capacity and grants, and a medium amount for inference to help bridge for portcos.
To be blunt, this is a very small amount of compute for neolabs, and only helps with bridging shortages, not with long term needs or ramps. But it hopefully provides a sense of the minimum sizes that will be needed to play in this space.
That's one example. More aggressive VCs will reserve larger clusters and count on follow-on funds to pay the remaining term, or explore alternative vehicles or liquidity lines to pay for the largest cluster they can get their hands on.
Finding, pricing, and closing on compute is challenging today and will get a lot harder in Q1. I'll have a longer blog post about this soon.
Chips & Chips ♠️
We’re hosting an invite-only poker night in San Francisco.
Thursday, august 27
7–10pm
food, cards, and a room full of builders.
space is limited.
link in thread
I asked people: "which AI supply chain startups (compute, power, chips, datacenters) valued at $10b or less are most interesting?"
10 votes: MatX
8: American Terawatt
6 each: Lightmatter, Modal, Prime Intellect
4 each: Valar Atomics, Wafer
3 each: Ayar Labs, d-Matrix, Gimlet Labs, Heron Power, Olix, Panthalassa, Positron, Substrate, Taalas, Tenstorrent, xLight
2 each: Aalo Atomics, Corintis, Fab2, General Matter, Ornn, SF Compute, Starcloud, Stone Power, TensorWave
1 each: 49 more startups, see image and below
People also voted for Base Power, Crusoe, Etched, and Fluidstack but those are over the valuation limit.
Disclosure: I'm a small angel in American Terawatt and Wafer. I didn't vote.
A lot of people ask me: Shouldn’t token prices be closely related to GPU prices? After all, GPU compute is the input and tokens are the output.
The answer is yes — in a perfectly competitive and operationally efficient market.
But we’re not there yet.
Running GPUs is one capability. Turning those GPUs into a highly efficient token factory — serving large models at scale with high utilization, low latency, optimized batching, routing, caching, and inference infrastructure — is another capability entirely.
We now have 300+ neoclouds, but only a handful of truly scaled token factories.
That gap matters. Token prices don’t just reflect the underlying cost of GPUs. They also reflect how efficiently compute is converted into intelligence.
Over time, as inference infrastructure matures and competition increases, I expect token economics and GPU economics to become much more tightly connected. But today, the conversion layer itself still carries significant scarcity value.
I remember pitching AMD to pms from top SMs and MMs about the cuda moat eroding back in November. Not only did no one believe me, no one cared.
Always found it funny that wall st refuses to genuinely learn the underlying technology and would rather focus 90% of their time on business cycles and lead times.
ROCm's improvement in the past 6 months has been exponential and this moves the needle drastically more than Helios landing a few weeks late.
Speed is the Moat 🚨🚨🚨 & @AnushElangovan & his team Keep Running Faster 🚀 & Faster! His team wants to run even faster, but one of the biggest blockers he hears about over and over again from his team is the lack of stable internal dev clusters.