Four innocent lines of PyTorch hide a video preprocessing footgun. In our H100 benchmark, the resulting physical plan made the model’s first operator run 8.9× slower. The code is clean. The physical plan isn’t. 🧵

Aug 13, 2026 · 6:36 PM UTC

1
2
13
2,261
With a CUDA VideoDecoder, the input pipeline is four lines. Every line seems reasonable.
1
3
123
But the tensor APIs also fix the physical plan: NVDEC → NV12 → RGB8 → resize + FP16 + normalize → RGB FP16 → first model operator The decoder and preprocessing each materialize an RGB tensor for the next stage.
1
3
86
At 4K, the normalized RGB FP16 tensor is 47.5 MiB. Preprocessing writes it. The first model operator reads it again. That is a 94.9 MiB boundary.
1
2
70
We ran the 4K, 3→8 first operator three ways on an NVIDIA H100 80GB HBM3: composed: 0.499 ms same conv, split: 0.107 ms same conv, fused: 0.056 ms The complete physical plans are 8.9× apart.
1
2
66
Fusion doesn’t always win though. At 64 output channels, the matched benefit falls to 1.2×: the fused schedule repeats preprocessing, while materialized RGB provides reuse.
1
3
80
Write the pipeline you want. Let Spiral choose the physical plan. Materialize RGB, keep native YUV, or fuse preprocessing into the model. Let the optimizer decide. See how @SpiralDB avoids the 9× footgun: spiraldb.com/blog/when-video…
1
4
84
Sort replies: Relevant Recent Liked