John Hahn retweeted
you'll know more about how KV caching, batching, and GPU communication affect transformer inference latency than 95% of people if you fully understand this article this is only the sixth resource in the ai performance engineering repo btw. follow and save to keep up with the series
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
10
24
405
37,533
so many hidden gems. make sure to check out the auxillary resources
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
4
329
John Hahn retweeted
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
19
19
180
46,772
John Hahn retweeted
ts compute trading
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 馃У
4
7
40
4,185
John Hahn retweeted
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 馃У
13
7
52
8,560
John Hahn retweeted
realtime inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
8
5
40
10,059
John Hahn retweeted
ultra fast inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
3
5
15
1,234
John Hahn retweeted
ultra fast inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
7
6
60
4,931
realtime inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
4
8
568
John Hahn retweeted
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
31
12
106
30,360
such a pleasure working with the @brilliantorg team!!
Ever since we moved our tutor to Wafer, the responses feel near-instant. We didn鈥檛 expect things to get this fast so soon. Thanks to @gpuemi and the @wafer_ai team! Love using a product built by a founder who grew up on Brilliant :)
1
3
316
John Hahn retweeted
Ever since we moved our tutor to Wafer, the responses feel near-instant. We didn鈥檛 expect things to get this fast so soon. Thanks to @gpuemi and the @wafer_ai team! Love using a product built by a founder who grew up on Brilliant :)
i grew up in Mexico, where it wasn鈥檛 always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn鈥檛 have to wait. on our dedicated GLM-5.2 endpoint, they鈥檙e now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it鈥檚 hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 馃У Brilliant鈥檚 full story in thread.
6
16
3,497
John Hahn retweeted
Replying to @gokulr @wafer_ai
馃憖
5
11
1,106
John Hahn retweeted
Replying to @gpuemi
It's been such a game-changer. We used to spend a lot of effort making the ~5s response times feel fast -- multiple layers of caching, loading animations, restructuring our prompts, etc. Switching to Wafer meant we could spend that time on the core parts of the learning experience instead!
5
11
455
John Hahn retweeted
A great learning experience has a sense of flow. Having a tutor take 5-6s to respond kills that. Switching to @wafer_ai for inference made Brilliant鈥檚 tutoring sessions feel fluid and fun again, and every session metric improved along with that.
i grew up in Mexico, where it wasn鈥檛 always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn鈥檛 have to wait. on our dedicated GLM-5.2 endpoint, they鈥檙e now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it鈥檚 hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 馃У Brilliant鈥檚 full story in thread.
8
5
30
2,744
John Hahn retweeted
@brilliantorg cut inference costs by 50% by making its AI tutor fast enough to answer in real time instead of pre-generating answers. to keep Koji, their AI tutor, responsive during math and coding lessons, the Brilliant team was speculatively prefetching ai-generated responses so students wouldn鈥檛 have to wait. on @wafer_ai 's dedicated glm-5.2 endpoint, Brilliant reached ~250 ms ttft and 300+ output tok/s. that speed let the team remove the speculative prefetch layer and have koji respond on demand. read how brilliant made the change. 馃У link in the thread.
2
4
48
1,684
John Hahn retweeted
i grew up in Mexico, where it wasn鈥檛 always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn鈥檛 have to wait. on our dedicated GLM-5.2 endpoint, they鈥檙e now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it鈥檚 hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 馃У Brilliant鈥檚 full story in thread.
13
14
92
19,982
John Hahn retweeted
@brilliantorg 3x鈥檇 their tokens per second and cut inference costs by 50%, by making its AI tutor fast enough to answer in real time instead of speculatively prefetching answers. To keep Koji, their AI tutor, responsive during math and coding lessons, the Brilliant team was speculatively prefetching AI-generated responses so students wouldn鈥檛 have to wait. On Wafer's dedicated GLM-5.2 endpoint, Brilliant reached ~250 ms time to first token and 300+ output tokens per second. This was a 3x throughput increase over their previous provider. That speed let Koji answer learners in real time, so the Brilliant team could stop pre-generating responses and remove the prefetching layer. Read how Brilliant made the switch! 馃У Link in Thread
9
7
28
1,335
John Hahn retweeted
you'll know more about compute and memory bandwidth bottlenecks than 95% of people if you fully understand this paper this is only the fifth resource in the ai performance engineering repo btw. imagine the ball knowledge in the other ones
we launched the most comprehensive ai performance engineering repo in the world now we'll be posting every single resource follow and save to keep up with the series. links in thread 馃У part 5: Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures the authors introduced a model for relating arithmetic throughput, memory bandwidth, and data reuse. applied to AI workloads, it gives a way to evaluate batching and kernel optimizations against the hardware's limits. the paper's concepts connect to these AI performance decisions: operational intensity measures floating-point work per byte of DRAM traffic after cache reuse. during dense transformer decoding, a linear layer processing a small token batch can read a large weight matrix for little arithmetic. compute and bandwidth roofs bound throughput. if weight reads limit an inference kernel, reducing transferred bytes can lower its memory-time bound. increasing peak arithmetic throughput alone leaves that bound unchanged. the ridge point marks the minimum intensity needed for peak compute throughput to be possible. batching tokens that share a weight matrix can increase work per weight byte, moving the matrix multiplication toward that threshold. computation ceilings account for limits in instruction parallelism and operation mix. for neural-network matrix multiplication, check Tensor Core use and compare throughput with the roof for the kernel's execution path and precision. bandwidth ceilings account for memory access patterns and data placement. for GPU tensor operations, changing the layout or thread-to-data mapping to coalesce scattered accesses can improve useful bandwidth. data reuse raises intensity when it reduces DRAM traffic for the same arithmetic work. a tiled matrix multiplication can reuse operands on chip; measure whether that reuse reduces DRAM byte traffic. the memory level determines which bytes to count. if a matrix multiplication reuses operands from L2, compare its L2 traffic with L2 bandwidth and its DRAM traffic with DRAM bandwidth. the authors demonstrated the model with 4 floating-point kernels on 4 multicore systems. for AI performance work, use it to assess weight reuse, memory layout, and arithmetic execution before choosing an optimization to test.
8
10
115
10,372