Helion × 🤗 Kernels is now live.
The key idea: autotuned, portable kernels can be packaged on the Hub with their pre tuned configurations, then loaded by users without local tuning or dependency friction.
Reported results:
• Helion attention beats PyTorch SDPA on all 19 tuned H100 shapes: 1.20× geomean speedup
• It wins on 9 of 10 held out shapes: 1.17× geomean speedup
• Seven linear attention variants beat FLA on all tuned B200 shapes: 1.41× on device and 1.33× end to end
• Forward plus backward: 1.55× geomean speedup
This makes the artifact being shared more than source code. It includes the performance decisions required to use the kernel efficiently.
pytorch.org/blog/helion-x-%F…
Sep 15, 2026 · 4:01 PM UTC
3
7
30
4,441


