we launched the most comprehensive ai performance engineering repo in the world
follow and save to keep up with the series. links in thread 馃У
part 6: Transformer Inference Arithmetic
Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s.
Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation:
- prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase.
- batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths.
- memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency.
- KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate.
- precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform.
- tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency.
- kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor.
- check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target.
Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.