365 days of GPU programming ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓░░░░░░░░░░░░ 264/365

out of memory
Pinned Tweet
Day 83/365 of GPU Programming Looking at DeepSeek's Multi-Head Latent Attention today. The last part of the AMD challenge series is to optimize an MLA decode kernel for MI355X where the absorbed Q and compressed KV cache are given and your task is to do the attention computation. A resource that really helped internalize what MLA does was @rasbt's incredible visual guide to attention variants in LLMs (luckily he posted that last week!), which covers everything from MHA to GQA to MLA to SWA, et cetera. If there's one place to get a visual intuition for recent attention mechanisms, it's this blog post. @jbhuang0604's video on MQA, GQA,MLA and DSA was the best conceptual intro I found on the topic and progressively builds up the ideas from first principles. The Welch Labs analysis of MLA is a great watch as well. Beautiful visualization of the changes DeepSeek made for MLA. Tried out a few kernels once I had a basic understanding of MLA and I think I'm slowly getting more comfortable with at least analyzing kernels.
Day 82/365 of GPU Programming Taking a closer look at Mixture of Experts today, so I can write better MoE kernels. Specifically, to optimize an MXFP4 MoE fused kernel for the GPU Mode challenge. I haven't had much prior exposure to MoEs, so lots of new concepts I learned today. Luckily I found the best intro to MoEs thanks to @MaartenGr visual overview of the topic. I then watched @tatsu_hashimoto's amazing Stanford CS336 lecture on MoEs, which added deeper context around why MoEs are gaining popularity, FLOPs, OLMoE, infra complexity, routing functions (mindblown this works so well...), expert sizes, training objectives, top k routing and DeepSeek variations. Once I had a basic understanding I started playing around with the some AITER kernels but progress there is tbd. Also had a nice chat with @juscallmevyom (who was kind enough to reach out!) about the AMD kernels and the challenge of materialization overhead.
23
145
1,457
129,969
Day 267/365 of GPU Programming Back to the Efficient AI Computing & TinyML course today. Going through lecture 5 on quantization and checking out the first assignments/labs of the course. I do have some basic familiarity with quantization and numeric data types (from learning about NVFP4/MXFP4) but excited to go a bit deeper on the topic.
Day 266/365 of GPU Programming Continuing my Triton explorations from yesterday and learning more about Gluon today. Its raison d'etre, how it fits into existing Triton code and just looking at examples and documentation. Probably won't go too deep into it but I want to get at least a high level understanding of the current state of Gluon.
1
3
31
922
levi retweeted
Have been busy with work, but finally I find the time to finish the blogpost on @GPU_MODE @coreauto's QR competition that happened 3 months ago. I didn't win the 1st prize, but my panel QR kernel is definitely the fastest (by a large margin) gau-nernst.github.io/b200-qr…
2
30
191
8,620
Day 266/365 of GPU Programming Continuing my Triton explorations from yesterday and learning more about Gluon today. Its raison d'etre, how it fits into existing Triton code and just looking at examples and documentation. Probably won't go too deep into it but I want to get at least a high level understanding of the current state of Gluon.
Day 265/365 of GPU Programming The recordings from last year's Triton developer conference showed up on my feed, so have been spending some time today going through some of the talks from the event. Outside of the AMD challenge, I haven't used Triton that much yet but a lot of the lessons shared by the speakers are applicable regardless of what you write your kernels with. For example, i really enjoyed the Performance Engineer's guide Blackwell GPUs by Chris Sullivan (a principal engineer at Nvidia).
2
3
45
2,649
Day 265/365 of GPU Programming The recordings from last year's Triton developer conference showed up on my feed, so have been spending some time today going through some of the talks from the event. Outside of the AMD challenge, I haven't used Triton that much yet but a lot of the lessons shared by the speakers are applicable regardless of what you write your kernels with. For example, i really enjoyed the Performance Engineer's guide Blackwell GPUs by Chris Sullivan (a principal engineer at Nvidia).
Day 264/365 of GPU Programming If you're interested in Nvidia's Tensor Cores, I think the guest lecture that @kimbochen gave at Cornell and Cohere Labs last year is quite an informative delineation of how they evolved over different GPU generations starting with Volt. The lecture itself is pretty compact but I enjoyed listening to the discussion around Hopper/Blackwell at the end of each lecture.
4
6
85
4,885
Day 264/365 of GPU Programming If you're interested in Nvidia's Tensor Cores, I think the guest lecture that @kimbochen gave at Cornell and Cohere Labs last year is quite an informative delineation of how they evolved over different GPU generations starting with Volt. The lecture itself is pretty compact but I enjoyed listening to the discussion around Hopper/Blackwell at the end of each lecture.
Day 263/365 of GPU Programming The folks at @togethercompute got access to Vera Rubin NVL72 and shared how they adapted their NVFP4 + FP8 GEMM kernels in ThunderKittens to bring them to cuBLAS comparable performance. Really recommend reading their blog post on their explorations even if you don't use TK! It's a great overview of what's to come in Vera Rubin and the implications and the implications for performant kernels: - Tensor Cores take 2x K - Tensor Memory increases to 576 columns - Shared Memory increases to 328 KiB - Collector Buffer functionality extends to B tile
1
5
48
4,659
Day 263/365 of GPU Programming The folks at @togethercompute got access to Vera Rubin NVL72 and shared how they adapted their NVFP4 + FP8 GEMM kernels in ThunderKittens to bring them to cuBLAS comparable performance. Really recommend reading their blog post on their explorations even if you don't use TK! It's a great overview of what's to come in Vera Rubin and the implications and the implications for performant kernels: - Tensor Cores take 2x K - Tensor Memory increases to 576 columns - Shared Memory increases to 328 KiB - Collector Buffer functionality extends to B tile
Day 262/365 of GPU Programming Going through lecture 4 of the Efficient AI Computing course which was livestreamed an hour ago. It's a continuation of the third lecture focusing on cost minimization/efficiency maximization via pruning and sparsity given a fixed compute budget. I'm especially interested in the section on system and hardware support for M:N sparsity especially interesting,
2
4
56
3,762
After this, people kept asking us how @modal is able to serve agent inference so well. So we wrote it all down. New blog: "How to serve trillions of tokens for trillion-parameter coding agents". modal.com/blog/trillion-toke…
found out we're serving 40% of Kimi K3 traffic on OpenRouter ... the guy working on it is tired. anyone wanna help us work on inference?
15
35
452
45,756
Day 262/365 of GPU Programming Going through lecture 4 of the Efficient AI Computing course which was livestreamed an hour ago. It's a continuation of the third lecture focusing on cost minimization/efficiency maximization via pruning and sparsity given a fixed compute budget. I'm especially interested in the section on system and hardware support for M:N sparsity especially interesting,
Day 261/365 of GPU Programming One thing I really appreciate about Simon's CUDA matmul kernel worklog is how clear each of his diagrams are. You can tell he put thought into how the components connect to one another such that there's a logical order to how you read the diagrams. I think a common problem with a lot of the GPU programming related visualizations out there these days (especially ones that are one shotted by AI) is that they are either too convoluted to follow or too simplified and underexplained. If you haven't given Simon's blog post a read yet, definitely set aside some time for it!
1
5
64
4,255
Day 261/365 of GPU Programming One thing I really appreciate about Simon's CUDA matmul kernel worklog is how clear each of his diagrams are. You can tell he put thought into how the components connect to one another such that there's a logical order to how you read the diagrams. I think a common problem with a lot of the GPU programming related visualizations out there these days (especially ones that are one shotted by AI) is that they are either too convoluted to follow or too simplified and underexplained. If you haven't given Simon's blog post a read yet, definitely set aside some time for it!
Day 260/365 of GPU Programming Started auditing a new course today (will try to cap it to 3 courses this fall)! While I'm waiting for new lectures and assignments to be released for MIT's Efficient AI Computing and Stanford's Deep Learning Alchemy courses, I will go through CMU's Inference Algorithms for Language Modeling from last fall. Inference (and especially inference engines) is a topic I want to understand much better, so this should be a nice complement to studying vLLM and SGLang. Also reading the rest of Simon's blog post and writing up some notes today.
2
8
92
5,188
Day 260/365 of GPU Programming Started auditing a new course today (will try to cap it to 3 courses this fall)! While I'm waiting for new lectures and assignments to be released for MIT's Efficient AI Computing and Stanford's Deep Learning Alchemy courses, I will go through CMU's Inference Algorithms for Language Modeling from last fall. Inference (and especially inference engines) is a topic I want to understand much better, so this should be a nice complement to studying vLLM and SGLang. Also reading the rest of Simon's blog post and writing up some notes today.
Day 259/365 of GPU Programming A blog post that's been on my to-read list for the longest time is Simon Boehm's How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog. Outside of PMPP and GPU Mode, it's probably the resource I've been recommended the most for studying GPU kernels. Will spend this weekend finally giving it the detailed read it deserves. Since the post is from 2022, my expectation is less about learning the most up-to-date optimization techniques and more so about understanding how a great performance engineer thinks through the process of optimizing a kernel step by step.
2
9
107
7,514
Day 259/365 of GPU Programming A blog post that's been on my to-read list for the longest time is Simon Boehm's How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog. Outside of PMPP and GPU Mode, it's probably the resource I've been recommended the most for studying GPU kernels. Will spend this weekend finally giving it the detailed read it deserves. Since the post is from 2022, my expectation is less about learning the most up-to-date optimization techniques and more so about understanding how a great performance engineer thinks through the process of optimizing a kernel step by step.
Day 258/365 of GPU Programming Continuing with the Efficient AI Computing course today. Lecture 3 kicks off the efficient inference chapter of the course which is what I'm most interested in. This lecture covered hardware support for sparsity, pruning granularity, N:M sparsity, scaling factors. Sparsity is a topic I generally want to understand better, particularly in relation to hardware codesign. Also grateful resources like this exist publicly and are released basically the same time as they're taught live at MIT. Excited to check out the assignments in Colab next!
2
16
178
13,120