Window for prompt speed: what turboderp's EXL3 packs change on a 32 GB card.
Same RTX 5090 (power limit 450 W), same TensorFold 0.6.2 (
@ashxhart; 0.6.3 is out), same Qwen3.8-27B + z-lab DFlash2 drafter, greedy. nvidia's NVFP4 checkpoint next to turboderp's EXL3 packs at 2.50bpw and 2.00bpw.
Window = max prompt+reply per request, set at boot (NVFP4 · EXL3-2.5 · EXL3-2.0, tokens):
parallel 1: 84,542 · 202,542 · 224,170
parallel 4: 29,180 · 143,085 · 163,959
parallel 8: 3,453 · 112,691 · 132,723
parallel 16 (patched build only): refused at boot (window 0) · 25,599 · 44,122
Drafted decode, one request, server at parallel 8, Ash's bench_concurrent with 256 tokens (median of 3 boots, 3 runs each): code prompt NVFP4 293.8 tok/s, EXL3-2.5 273.3 tok/s; chat prompt 204.0 tok/s, 138.7 tok/s. Cold prefill on a 2k prompt, same servers (median of 18 runs): 8,031 tok/s, 2,543 tok/s.
EXL3-2.0, same servers, 2 boots, one patched: decode 362.1 tok/s code, 209.3 tok/s chat (median of 2 boots, 3 runs each), prefill 2,540 tok/s (median of 12 runs).
So both EXL3 packs pay for their window in prompt speed. Decode is mixed: EXL3-2.0 was faster than NVFP4 on both prompts, EXL3-2.5 slower.
Token-exact: drafts matched serial in all 16 suite arms. Needle found in all 16 runs (31,533-token prompts on the EXL3 arms, 1,008-token prompts on NVFP4, because at parallel 8 its window is 3,453).
Honest limits: I measured window and speed, not quality. The packs sit far apart in bits per weight (NVFP4 ~4.5 effective on its FP4 layers, FP8 and bf16 around them; EXL3 2.50 and 2.00), so more window does not mean the same answer quality. TensorFold boots EXL3 packs as experimental. One run day, one card.
My 16 GB tight-staging patch (earlier post) ran in the patched arms: in the A B B A A B pairs it changed neither window nor speed (medians within 0.4 %). NVFP4 boot peak, highest of 6 servers each on the 1 s sampler: release 23.4 GiB, patched 20.9 GiB. The patch is not in TensorFold.