Agentically Writing a Fused Blockwise Hadamard Quantization Kernel in CuTe DSL
Quantization has become one of the most widely used optimizations in LLM inference, reducing both memory and computational costs, making it possible to run large models that wouldn’t fit in memory