Session 1 of the KubeSimplify Book Club is done. We're reading Inference Engineering by Philip Kiely and this week covered the Preface, Chapter 0, and Chapter 1.
Chapter 0:
- open models (Hugging Face, Nvidia) are closing the gap with closed frontier models, a big reason inference engineering is hot in 2026
- local inference means running models on your own hardware
- every decision comes down to latency, availability, and cost
- the pipeline: training, learning, inferencing, serving
on the ground: teams renting GPU racks from providers like NeevCloud, enthusiasts buying DGX Spark and Nvidia GPUs
Chapter 1:
- model weights are layer 1, inference runtime is layer 2
- speculative decoding to speed up generation
- the sumo wrestler vs swimmer analogy for throughput vs latency tradeoff
- metrics to know: TTFT, TPS, P90
- tensor parallelism to scale across GPUs
Recording is up, link below. Next session Friday 8:30 PM IST with Saiyam.