FLOPs are a poor proxy for kernel performance. Arithmetic intensity is closer to the real constraint.
Consider a GEMM:
C = A × B
A: M × K
B: K × N
C: M × N
The computation requires approximately: 2MNK FLOPs
But the hardware doesn’t execute FLOPs in isolation.
Data has to move through the memory hierarchy: HBM → L2 → SRAM/Shared Memory → Registers
For a naive implementation, the same elements of A and B may be fetched repeatedly.
A tiled kernel changes the problem:
Global Memory
↓
Tile Load
↓
Shared Memory
↓
Register Blocking
↓
FMA
The objective isn’t simply to reduce the number of instructions.
It is to increase data reuse per byte fetched from the expensive memory level.
That’s why two kernels with identical FLOP counts can have completely different performance.
A useful first-order model is:
AI = FLOPs / Bytes transferred
P ≤ min(P_peak, BW × AI)
But there is an important caveat:
The denominator is not necessarily the tensor size.
It is the amount of data actually transferred at the memory level that is limiting execution.
A kernel can therefore have high theoretical arithmetic intensity and still become memory-bound because of poor tiling, register spills, cache misses, synchronization, or insufficient occupancy.
At that point, adding more TFLOPS doesn’t solve the problem.
The optimization target becomes data movement and reuse, not computation.
This is also where GPU and NPU performance analysis starts becoming interesting: the same computational graph can produce very different bottlenecks after operator lowering, tiling, fusion, and runtime scheduling.
Performance is not determined by how much computation a model contains. It is determined by how efficiently the hardware can move and reuse the data required to perform that computation.
#MLSystems #PerformanceEngineering #GPU #NPU #AIInfrastructure
O(n²) tells you the scaling. It doesn’t tell you the runtime.
scores = Q @ K.transpose(-2, -1)
attn = softmax(scores)
output = attn @ V
With sequence length n and head dimension d, the attention matrix scales as:
QKᵀ → O(n²d)
The problem is that FLOPs are only part of the story.
For long sequences, the intermediate scores tensor can become a serious memory and bandwidth cost:
scores.shape = [batch, heads, n, n]
At n = 16K, h = 32, even FP16 requires roughly 16 GB just for the attention matrix of a single batch.
This is one of the reasons memory-efficient attention implementations don’t simply try to “compute faster”.
They change how the computation is tiled and how much intermediate data ever reaches HBM.
Algorithmic complexity tells you what scales.
The memory hierarchy often tells you what actually runs fast.
#MLSystems #DeepLearning #Attention #Inference
Kernel fusion can make a kernel faster and the model slower.
I’ve seen this happen when the fused kernel reduces launch overhead but increases register pressure and spills intermediate data to local memory.
The kernel benchmark looks better.
The end-to-end latency doesn’t.
That’s the part of optimization that’s easy to miss: changing the kernel also changes its resource profile — registers, shared memory, occupancy, memory traffic, and sometimes even the scheduling behavior of the runtime.
So I don’t treat kernel speedup as an end goal.
A kernel optimization is only meaningful if the workload gets faster on the target hardware.
#MLSystems #CUDA #Inference #PerformanceEngineering
Peak FLOPS is one of the least useful numbers when debugging inference performance.
What matters is how much work the hardware can actually keep in flight.
An operator with low arithmetic intensity can spend most of its time moving data rather than doing computation. More compute capacity won’t help if memory bandwidth is already the limiting factor.
And once you optimize that kernel, the bottleneck may simply move to synchronization, kernel launches, or host-device transfers.
That’s why I prefer looking at the execution trace and arithmetic intensity before drawing conclusions from GPU utilization or theoretical TFLOPS.
Hardware performance is a workload property, not a specification-sheet number.
#MLSystems #PerformanceEngineering #AcceleratedComputing
Banger report from Microsoft.
(bookmark it)
They show that it's possible to build competitive small coding agents without traditional distillation from frontier models.
This is a big deal!
The work describes how they achieved this.
They introduce a 4B coding agent trained on roughly 1,500 software engineering environments.
The cool thing is that they use no distillation from a larger model at any point.
FrogNano is post-trained purely with RL on synthetic tasks.
The target is a coding agent that runs on minimal machines, which rules out both a frontier backbone and a frontier teacher.
The ingredient the report credits the most is online task synthesis.
The pipeline generates tasks calibrated to the frontier of learnability for the current checkpoint, so the agent always trains on problems it can just barely solve. The authors argue that calibration, rather than the volume of synthetic data, is what makes this work.
This means that competitive small coding agents can be trained from synthetic tasks alone.
And generating those tasks at the current agent's learnability frontier is what makes this particular training productive.
The report covers training methodology, evaluations across diverse environments, and analyses of what the agent learned.
Paper: https://t.co/rSmH21XMD5
A model being fast in isolation doesn’t mean the inference system is fast.
End-to-end performance is often determined outside the model itself.
Model execution
→ Operator scheduling
→ Kernel efficiency
→ Memory movement
→ Host–Device synchronization
→ Runtime overhead
→ Request scheduling
→ Batching strategy
A kernel-level optimization may reduce compute time while increasing synchronization overhead.
Larger batches may improve GPU utilization while increasing tail latency.
Faster kernels may deliver little system-level gain if memory movement remains the bottleneck.
This is why inference optimization should be measured end-to-end, not just at the model or kernel level.
Optimize the critical path, not the isolated operation.
#MLSystems #Inference #AIInfrastructure
Inference performance is rarely limited by compute alone.
A model can have enough FLOPs, yet still fail to achieve high accelerator utilization.
The real bottleneck may be:
Compute-bound → insufficient arithmetic throughput
Memory-bound → bandwidth becomes the bottleneck
Kernel launch overhead → too many small operations
KV Cache → memory capacity and bandwidth dominate at scale
Batching → higher throughput, but potentially worse latency
This is why FLOPs ≠ Performance.
A serious inference optimization process should profile the entire execution path:
Model → Operator → Kernel → Memory → Runtime → Hardware
Don’t optimize the model in isolation.
Optimize the system.
#AI #Inference #MLSystems #AcceleratedComputing