Introducing VC-Attention: fast and accurate low-bit attention without retraining.
On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods.
Two key innovations:
• V-Smooth reduces value quantization error.
• ExpCast-FP8 speeds up softmax.
Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300.
Blog:
nunchux.ai/blog/attention-is…
Technical Report:
arxiv.org/pdf/2609.15810
Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.