Just finished Inference Engineering book by Philip Kiely.
If you are shipping LLMs behind an API, this is the map. Not the hype, the actual math on batching, KV cache, and why your GPU is at 40%.
30 lessons over the next month. Thread below ๐
If you're learning inference engineering today, you're not late.
The field is 3 years old. The best practices are still being written.
Kiely's book is the first comprehensive map. The territory is still expanding.
Build things. Measure everything. Share what you learn.
#END
30/30-from Inference Engineering (IE) by Philip Kiely (The Future of Inference Engineering)
The future of inference engineering: We are moving from "run the model" to "orchestrate the compute."
Disaggregated inference: Prefill on one GPU, decode on another.
The inference engineer of 2027 will spend less time on CUDA kernels and more time on:
* Cluster topology
* Request routing algorithms
* Hardware heterogeneity
* Energy efficiency
It's becoming a systems problem, not a GPU programming problem.
Custom silicon will reshape cost curves.
Groq's LPU: Deterministic latency, no batching needed.
AWS Inferentia: 40% cheaper than a GPU for standard workloads.
Google TPU: Best price/performance for Google's models.
The NVIDIA monopoly is cracking.
The prefill/decode split is the big trend.
Prefill needs compute. Decode needs memory bandwidth.
They are different hardware profiles.
Soon: Prefill clusters (H100) + Decode clusters (L4) + Router in between.
Specialization beats generalization.
Need to serve 70B+ models? --> A100/H100 with TP Need to serve many small models? โ Multi-GPU node, one model per GPU
Match hardware to workload. Don't default to "the best."
29/30-from Inference Engineering (IE) by Philip Kiely
Hardware selection decision tree:
Need <50ms latency? --> L4/T4 edge
Need <200ms, high throughput? --> A100
Need <100ms, highest throughput? --> H100
When in doubt, add observability.
Log:
* Input/output token counts
* Generation time per token
* GPU memory per layer
* CUDA kernel execution times
You can't optimize what you can't see.
28/30-from Inference Engineering (IE) by Philip Kiely (Production Debugging)
Debugging inference in production:
"It works on my machine" is bad.
"It works at batch size 1" is worse.
Shadow traffic is your friend.
Send 1% of production traffic to the new version. Compare outputs, latency, and memory.
Don't trust unit tests for inference. Trust statistical comparisons.
Reproduction checklist:
* Same model weights (checksum them)
* Same framework version
* Same CUDA/driver version
* Same batch size and sequence length
* Same quantization config
Any mismatch invalidates the test.
Production bugs that only appear at scale:
* OOM at batch size > 4
* Race conditions in continuous batching
* NaN outputs after 2K tokens
* Memory leaks in KV cache
27/30-from Inference Engineering (IE) by Philip Kiely
The best inference optimization is often not computing at all.
Cache layers:
* Exact prompt match โ return cached response
* Similar prompt โ semantic cache (embeddings)
* Prefix match โ reuse KV cache
Smaller images = faster cold starts, faster CI, faster deploys.
In inference engineering, deployment speed matters as much as inference speed.
Every second of downtime is tokens not served.
26/30-from Inference Engineering (IE) by Philip Kiely (Containerization)
Your inference Docker image is probably 10GB too big.
Base PyTorch image: 4GB+
Model weights: 14GB+
CUDA drivers: 2GB+
You are shipping 20GB for a 100MB application.
Slimming strategies:
Use NVIDIA's CUDA runtime base, not devel
Multi-stage builds: compile in one image, copy binary to another
Exclude training dependencies (transformers[training] is huge)
Use Safetensors instead of pickles
Target: <5GB for most LLM serving images.
25/30-from Inference Engineering (IE) by Philip Kiely
Multi-tenancy in inference:
Your GPU is a timeshare, not a private island.
One user shouldn't monopolize VRAM because they sent a 32K context.
Set limits:
* Max tokens per request
* Max context length per tenant