Filter
Exclude
Time range
-
Minimum likes
Replying to @gpusteve

ALT Ben Stiller Knowledge GIF

2
84
John Hahn retweeted
you'll know more about how KV caching, batching, and GPU communication affect transformer inference latency than 95% of people if you fully understand this article this is only the sixth resource in the ai performance engineering repo btw. follow and save to keep up with the series
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
10
24
407
37,671
so many hidden gems. make sure to check out the auxillary resources
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
4
329
Replying to @wafer_ai
holy knowledge
1
3
179
John Hahn retweeted
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 馃У part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
19
19
180
47,135
Replying to @gpusteve
trading ts compute
34
John Hahn retweeted
ts compute trading
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 馃У
4
7
40
4,195
Replying to @wafer_ai @gpugene
ts compute trading
2
74
John Hahn retweeted
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 馃У
13
7
52
8,589
Replying to @gpugene
eugenius strikes again
3
580
Replying to @gpuemi

ALT Driving Richard Hammond GIF by DriveTribe

1
30
John Hahn retweeted
realtime inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
8
5
40
10,065
Replying to @gpubiz
we're kinda tuff ngl
1
14
John Hahn retweeted
ultra fast inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
3
5
15
1,234
Replying to @gpusteve
:o
1
57
John Hahn retweeted
ultra fast inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
7
6
60
4,932
realtime inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
4
8
568

ALT The Incredibles Running GIF by Disney+

1
2
356
John Hahn retweeted
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC鈥檚 request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 馃У link in thread
31
12
106
30,384