we gathered all the resources you'll ever need to become the most cracked ai performance engineer
follow and save to keep up with the series. links in thread 🧵
part 7: Efficiently Scaling Transformer Inference
Reiner Pope and coauthors studied PaLM inference on TPU v4. the paper gives performance engineers a way to reason about where weights, activations, and KV caches should live, what must move each step, and which transfers limit latency.
the paper's mechanisms suggest these checks for an inference deployment:
- prefill and decode can have different arithmetic intensity, meaning computation per byte transferred. prefill can process many prompt tokens in one pass; standard autoregressive decode produces one new token per sequence per step. larger decode batches amortize weight reads, but add sequence-specific KV cache traffic. sweep batch size with context length and profile the linear layers and attention to locate the bandwidth or compute limit.
- tensor parallelism can reduce local compute without removing the communication bottleneck. in the paper's 1D layout, activation-aggregation time stays roughly constant as chip count grows for a fixed workload. 2D partitioning splits both weight dimensions so each chip communicates smaller activation shards. evaluate this against the actual interconnect and matrix dimensions before increasing the parallelism degree.
- the best tensor to move changes with the token batch. weight-stationary layouts exchange activations; weight-gathered layouts transfer more weights to reduce activation exchange. large prefills can justify that trade because their activation tensors are large. count the tokens processed per forward pass, including prompt tokens during prefill, when comparing those layouts.
- multiquery attention shares one KV head across query heads, but a head-sharded layout can replicate that cache across devices. the paper shards attention across the batch during decode and exchanges current-step activations to reduce cache reads per chip. in vLLM's grouped-query attention deployments, check whether tensor parallelism exceeds the KV-head count and introduces cache replication before estimating per-GPU capacity.
- KV cache capacity and prefill attention intermediates are separate memory problems. at a fixed batch size and query-head count, sharing KV heads does not shrink the full attention-score tensor. for dense prefill, that tensor's element count grows with the square of sequence length. Appendix G discusses microbatching to reduce temporary allocations and FlashAttention to avoid materializing the full score matrix. use that distinction to identify whether a memory failure comes from retained KV state or temporary attention allocations.
- communication overlap depends on the execution schedule and tensor layout. the authors use Looped CollectiveEinsum to overlap collectives with matrix multiplication, including choosing output sharding that exposes overlap opportunities. inspect a GPU timeline for exposed communication and dependent compute stalls; an asynchronous collective does not establish that its cost is hidden.
- quantization changes different bottlenecks depending on what is quantized. the paper stores weights in int8 while retaining bfloat16 matrix arithmetic. this reduces weight-loading traffic, but does not provide faster arithmetic for compute-bound batches. evaluate weight, activation, and KV cache precision against the measured bottleneck and benchmark the effect on both phases.
- benchmark the operating point that the application needs. the paper compares latency with accelerator-time per token and model FLOPS utilization. in a serving system, track time to first token, inter-token latency, and throughput under the target load. keep queueing time separate from prefill time.
the practical use is to narrow an optimization experiment: estimate compute, HBM traffic, and collective costs for the actual workload, inspect the bottleneck, change the relevant layout or precision, and remeasure at the same latency target.
Andrej Karpathy spent 8 years at OpenAI and Tesla
Last week, he condensed everything he knows into one free 2-hour lecture
Agents → Loops → Graphs → Self-Improving Systems
People pay $25K for bootcamps that teach less than this
This lecture beats most paid AI engineering courses
You probably don't have 2 hours right now
Don't let this disappear from your feed
Bookmark and watch it
Then read the article below
This is huge.
Xiaomi just open-sourced 7k+ reinforcement learning task environments used to train MiMo, covering code, cybersecurity, general tool use, visual web development, and music-related tasks.
here's why this matters more than another model release:
open weights let you run a model. open environments let you teach one.
reinforcement learning environments are difficult to build, because each task needs an executable world, a clear objective, tools the agent can use, and a reliable test for success.
Xiaomi has released much of that missing infrastructure, including prompts, agent configurations, reward metadata, and the surrounding training framework.
here is what developers can now do with it.
→ train models on real agent work. instead of imitating static answers, a model can attempt tasks, call tools, observe results, and improve from the outcome.
→ specialize smaller models. teams can train an open checkpoint for software engineering, vulnerability reproduction, knowledge work, or web development instead of relying on a much larger general-purpose model.
→ study complete trajectories. each run captures the agent’s reasoning, tool calls, environment feedback, and reward. researchers can identify where agents fail and compare training algorithms, reward strategies, harnesses, and models.
→ build a complete training loop. Xiaomi’s stack covers environment interaction, trajectory collection, reward evaluation, and policy optimization. developers can replace individual components and measure what improves behavior.
the key shift is from open-sourcing model outputs to open-sourcing learning experiences.
a dataset teaches a model what an answer looks like. an environment lets it discover which actions lead to success.
check this out on HuggingFace 🤗: https://t.co/BiULCVsIMB
i wrote a complete breakdown of fine-tuning models with reinforcement learning using GRPO and RULER. RULER uses a language model judge to rank trajectories, so you do not have to manually write reward functions or collect labeled answers.
the article is quoted below.
Reasoning from scratch, round number 5! This time, talking about log-probability scoring (also a great fundamental concept for loss functions like cross-entropy in pre-training and distillation) and self-refinement.
00:00 Introduction and inference-time scaling recap
05:02 Loading the pretrained LLM
08:00 Comparing and scoring model answers
10:18 Building a rule-based scorer
17:53 Token probabilities and sequence likelihood
26:47 Computing token probabilities in PyTorch
30:12 Token indexing and shifted targets
37:27 Log probabilities and numerical stability
45:57 Scoring answers with average log probabilities
56:24 How self-refinement works
59:07 Generating critiques and revised answers
1:01:00 Implementing the self-refinement loop
1:05:57 MATH-500 evaluation results
1:07:35 Takeaways and next steps
It’s easy to dismiss Jev it as “just a classifier”.
But people (me included) who have been training encoder-style models for classification for many years know they were usually special-purpose and limited in some way.
The breakthrough of Jev is that it generalizes well.
And I’d say the secret sauce is probably more in the data than in the training algorithm. (Plus a nice API design on top of it.)
Day 259/365 of GPU Programming
A blog post that's been on my to-read list for the longest time is Simon Boehm's How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog. Outside of PMPP and GPU Mode, it's probably the resource I've been recommended the most for studying GPU kernels.
Will spend this weekend finally giving it the detailed read it deserves. Since the post is from 2022, my expectation is less about learning the most up-to-date optimization techniques and more so about understanding how a great performance engineer thinks through the process of optimizing a kernel step by step.
you'll know more about CUDA than 90% of people if you fully understand this guide
this is only the third resource in the ai performance engineering repo btw.
imagine the ball knowledge in the other ones
Inference scaling part 1.
Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x)
00:00 Introduction and recap
00:31 Training-time and inference-time scaling
07:52 What we'll implement
11:47 Notebook setup and model loading
17:43 Building a flexible text generation function
24:40 Chain-of-thought prompting
28:26 Sampling and output diversity
33:43 Next-token logits and greedy decoding
38:20 Temperature scaling step by step
42:46 Softmax and token probabilities
47:42 Multinomial sampling
54:51 Adding temperature sampling to text generation
59:31 Top-p filtering step by step
1:10:23 Adding top-p filtering to text generation
1:13:43 Sampling and LLM watermarking
1:16:01 Self-consistency and majority voting
1:20:36 Implementing self-consistency
1:29:02 MATH-500 results
1:35:01 Accuracy and compute tradeoffs
1:36:50 Next steps and self-refinement
Google's Jeff Dean just released the best 1-hour lecture on AI engineering: from basics to Graphs
1:45 - LLM from scratch
17:22 - how to use AI models
30:03 - prompt engineering
52:35 - one human coordinating 100 agents
1:02:40 - where the coordination actually lives
27 years of building AI at Google, compressed into one hour
Prompts → Agents → Loops → Graphs
most people will stop at the prompt engineering chapter and call it learning
he spends the last twenty minutes on the part that is still true next year
same model, same tokens, completely different week
watch it today
the full guide on graph engineering is below, save it while it is still early ↓
Build your own harness, folks.
This is absolute banger paper from NVIDIA on self-evolving agent harnesses.
(bookmark it)
They introduce SoL-Pi which cuts token traffic by nearly half.
And it matches its baseline harness on GPT-5.6 Sol and Opus 5.
More details below:
Instead of tuning a harness by hand, they run auto-research loops at the harness layer across many repository-derived and verifier-driven environments, keeping only the mechanisms that survive selection.
Four mechanisms survived:
> Action Fusion changes how actions execute
> Online Context Compact handles compaction during a run
> ObservationPack reshapes observation handling
> Evidence-Preserving Reducer covers delegated reading
On the 51-task EdgeBench evaluation, the savings translate to about a third off API cost. In dollars that is an estimated $8.75 to $13.50 per hour against native Codex and Claude Code harnesses, and $4.36 to $5.71 against the baseline harness.
Because the search runs across many environments rather than one, the retained mechanisms keep working outside the setting that produced them. Code is on GitHub under NVlabs.
Paper: https://t.co/1x26LzuE6d
Chat with Paper: https://t.co/kygTc5XLFB
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
@rasbt@rasbt Love your work. Where do you even get time to get on top of all these changes? Awesome!!
Looking forward to your DeepSeek v4.1 blog. Cheers!!
"A Mathematical Explanation of Transformers" is a recent paper that develops a rigorous mathematical framework for understanding the architecture behind Transformers and large language models.
It interprets the Transformer as a discretization of a continuous integro-differential equation, with self-attention represented as a non-local integral operator, layer normalization as a projection onto a constrained set, and feedforward layers and activation functions incorporated into the same mathematical framework, then uses operator splitting and numerical discretization to recover the standard Transformer architecture and extend the formulation to multi-head attention, Vision Transformers, and convolutional Transformers.
I've already shared several resources on the mathematics behind neural networks, Transformers, and LLMs, but there always seems to be something new and interesting to explore in this area.
https://t.co/18IZwRRDZS
pretty cool new course by stanford this fall on “engineering ai agents”
goes through a lot of fundamentals of complex ai systems design
https://t.co/8mExYeck9h
this is f**king insane.
a solo dev just open sourced a 100% FREE ElevenLabs replacement that runs entirely on your own machine.
the GitHub repo is at 19.4K stars.
it lets you:
→ clone a voice from one clean reference clip
→ dub any video into 646 languages
→ generate audiobooks, dictation, transcription
→ pick from 14 TTS engines instead of one
ElevenLabs supports 32 languages. this does 646.
no per-character billing. no usage caps. no audio ever leaves your computer.
save this for later.
repo below
I’ve spent 10 years teaching math to machine learning engineers.
80% of university math is irrelevant to your actual job.
Here's the 20% you actually need to build models (and how to learn it fast):
https://t.co/sV52SBB16J