Helion × 🤗 Kernels is now live. The key idea: autotuned, portable kernels can be packaged on the Hub with their pre tuned configurations, then loaded by users without local tuning or dependency friction. Reported results: • Helion attention beats PyTorch SDPA on all 19 tuned H100 shapes: 1.20× geomean speedup • It wins on 9 of 10 held out shapes: 1.17× geomean speedup • Seven linear attention variants beat FLA on all tuned B200 shapes: 1.41× on device and 1.33× end to end • Forward plus backward: 1.55× geomean speedup This makes the artifact being shared more than source code. It includes the performance decisions required to use the kernel efficiently. pytorch.org/blog/helion-x-%F…

Sep 15, 2026 · 4:01 PM UTC

3
7
30
4,441
Sort replies: Relevant Recent Liked
Replying to @LysandreJik
love that the hub object now carries the performance decisions, not just the kernel text
36