Four innocent lines of PyTorch hide a video preprocessing footgun.
In our H100 benchmark, the resulting physical plan made the model’s first operator run 8.9× slower.
The code is clean. The physical plan isn’t. 🧵
Aug 13, 2026 · 6:36 PM UTC
1
2
13
2,261
But the tensor APIs also fix the physical plan:
NVDEC → NV12 → RGB8
→ resize + FP16 + normalize → RGB FP16
→ first model operator
The decoder and preprocessing each materialize an RGB tensor for the next stage.
1
3
86
At 4K, the normalized RGB FP16 tensor is 47.5 MiB.
Preprocessing writes it. The first model operator reads it again.
That is a 94.9 MiB boundary.
1
2
70
We ran the 4K, 3→8 first operator three ways on an NVIDIA H100 80GB HBM3:
composed: 0.499 ms
same conv, split: 0.107 ms
same conv, fused: 0.056 ms
The complete physical plans are 8.9× apart.
1
2
66
Fusion doesn’t always win though. At 64 output channels, the matched benefit falls to 1.2×: the fused schedule repeats preprocessing, while materialized RGB provides reuse.
1
3
80
Write the pipeline you want. Let Spiral choose the physical plan.
Materialize RGB, keep native YUV, or fuse preprocessing into the model. Let the optimizer decide.
See how @SpiralDB avoids the 9× footgun:
spiraldb.com/blog/when-video…
1
4
84

