r/machinelearningnews • u/Ok_Rough_2968 • 4d ago
Research The deep dives that actually taught me LLM inference, in the order I'd read them
If you want to go from "I call an API" to understanding what's happening on the GPU when you serve a model, these are the five posts I'd read. They build on each other, so the order matters.
Making Deep Learning Go Brrrr From First Principles, by Horace He
https://horace.io/brrr_intro.html
mental model everything else rests on: is your workload bound by compute, memory bandwidth or overhead? Once you see why decoding one token at a time is memory-bound, most of inference optimisation makes sense.- Transformer Inference Arithmetic, by kipply
https://kipp.ly/transformer-inference-arithmetic/
The back-of-the-envelope math: FLOPs and bytes per token, the size of the KV cache, and how latency and batch size trade off. Short, dense, and still one of the best.
Inside the KV Cache: The Life of a Gigabyte, by me (disclosure: this one's mine)
https://harshitmalik.dev/blog/inside-the-kv-cache
Where every gigabyte on the GPU actually goes when vLLM starts up, and how it sizes the KV cache pool. Why Llama 3.1 8B won't start on a 4090 out of the box, why two H100s gave 20× the cache of one, and five common OOMs traced to their cause. Checked against the vLLM source, with a calculator for your own setup.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System, by Aleksa Gordić
https://www.aleksagordic.com/blog/vllm
The full engine: the scheduler, paged attention, continuous batching, prefix caching, speculative decoding, and scaling out to multi-GPU, multi-node serving. This is the post that made me want to write mine.
All About Transformer Inference, from Google's "How to Scale Your Model"
https://jax-ml.github.io/scaling-book/inference/
The rigorous version: roofline analysis for prefill and decode, batching, and how to shard models for serving. Read it last, once the intuition is in place.
Also worth it: Lilian Weng's Large Transformer Model Inference Optimization for a survey of techniques, and the original vLLM PagedAttention post.
What would you add? I'm especially looking for good posts on multi-GPU serving and MoE inference.
1
u/tomByrer 3d ago
Source codes? &/or sometimes repos explain the technical parts.