# GLM-5.3-Flash (FP8) on 8x DGX Spark — TP8 serving info (reproducible)
## TL;DR
World-first (AFAIK) GLM-5.3-Flash on 8x DGX Spark (GB10/sm_121a), TP8, **official zai-org FP8 checkpoint**
(not NVFP4), 1M context, 4.39M-token KV pool, NEXTN MTP-5, vision working.
Built on
@0xSero's sm121 bundle — extended from TP4/NVFP4 to TP8/FP8 with two extra workarounds.
## Hardware
- 8x NVIDIA DGX Spark (GB10, sm_121a, 121.69 GiB unified memory each)
- RoCEv2 dual-rail mesh: 2x ConnectX-7 per node (200GbE effective/node), MikroTik CRS804
- 1 GPU per node → TP_SIZE = EP_SIZE = NNODES = 8
## Stack
- Base image: lmsysorg/sglang:glm-5.3-flash (day-0 GLM-5.3 image)
- Patches:
github.com/0xSero/glm-5.3-fl… (6 baked patches — glm5_next model,
NEXTN draft, sm120 quant utils, modelopt quant, flash_mla_sm120 w/ NoPE zero-rope path, dsa_backend)
- Rebuilt with TORCH_CUDA_ARCH_LIST=12.1a / CUTE_DSL_ARCH=sm_121a, linux/arm64
- Model: zai-org/GLM-5.3-Flash (native FP8, 306 GiB, 62 shards) — every node holds a full local copy
## Why vLLM does NOT work (as of 2026-08-27)
GLM-5.3-Flash MLA is NoPE (qk_rope_head_dim=0). The only attention backend vLLM offers on sm_121
(FLASHINFER_MLA_SPARSE_SM120) force-selects the fp8_ds_mla KV format, whose kernel hard-codes
pe_dim=64 → `concat_and_cache_mla: pe_dim must be 64 for fp8_ds_mla`, crash at ~55% load, 3/3 repro.
Not bypassable by --kv-cache-dtype (auto is overridden by the backend).
## Two extra workarounds needed for FP8 + TP8
1. **SGLang Fp8MoEMethod bug**: with `--moe-runner-backend flashinfer_cutlass` (the bundle default),
create_moe_runner() silently skips runner init for cutlass ("else: pass # TODO") →
`AttributeError: 'Fp8MoEMethod' object has no attribute 'runner'`.
Workaround: `--moe-runner-backend triton`.
2. **RDMA memory registration**: add `--cap-add IPC_LOCK --ulimit memlock=-1:-1` to docker run,
or NCCL dies with `ibv_reg_mr_iova2 failed: Cannot allocate memory`.
Also: `--enable-prefill-cp` (zigzag) crashes in cuda-graph capture (deep-ep invalid resource handle)
on this combo — do not use.
## Launch (per node, rank 0..7, workers use same cmd)
docker run -d --name glm53 --gpus all --network host --ipc host --shm-size 32gb \
--cap-add IPC_LOCK --ulimit memlock=-1:-1 --device /dev/infiniband:/dev/infiniband \
-v /var/tmp/models:/cache/huggingface \
-e NCCL_NET=IB -e NCCL_IB_HCA=<your_rocev2_hcas> -e NCCL_SOCKET_IFNAME=<your_ifaces> \
-e NCCL_IB_GID_INDEX=3 -e NCCL_MAX_NCHANNELS=8 -e NCCL_MIN_NCHANNELS=8 -e NCCL_BUFFSIZE=16777216 \
<sm121-patched-image> \
python3 -m sglang.launch_server \
--model-path /cache/huggingface/GLM-5.3-Flash-FP8 \
--served-model-name glm-5.3-flash --host 0.0.0.0 --port 8210 \
--tp-size 8 --ep-size 8 --nnodes 8 --node-rank
$RANK \
--dist-init-addr <head_mesh_ip>:27000 \
--context-length 1048576 \
--attention-backend dsa \
--dsa-prefill-backend flashinfer_sparse_mla --dsa-decode-backend flashinfer_sparse_mla \
--linear-attn-backend triton \
--kv-cache-dtype fp8_e4m3 \
--moe-runner-backend triton \
--disable-shared-experts-fusion \
--chunked-prefill-size 2048 --max-prefill-tokens 2048 \
--max-running-requests 8 --mem-fraction-static 0.83 \
--cuda-graph-max-bs-decode 8 \
--speculative-algorithm NEXTN --speculative-num-steps 5 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 6 --speculative-adaptive \
--trust-remote-code
(note: single value only for --served-model-name; NCCL ifaces are cluster-specific)
## Non-obvious tuning finds
- **--chunked-prefill-size 2048 doubled long prefill** vs the "bigger is better" intuition:
70K-token prefill = 887 tok/s
@8192 → **1,748 tok/s
@2048** (1024 regresses to 1,575).
Bonus: prefill window halves, so concurrent-decode starvation window halves too.
- mem-fraction-static 0.83, not 0.90: GB10 has a boot-time usable ceiling (~101.6/121.69 GiB).
- NCCL_MAX/MIN_NCHANNELS=8 (from our GLM-5.2 tuning — 2.6x prefill vs 4 channels on this mesh).
- cold first request ~232 tok/s (triton JIT) — warm up before judging numbers.
- Engine startup ~6 min (weights 235s + cuda-graph capture).
## Numbers (see screenshot)
prefill (cache-busted): 7K=1,842 / 37K=1,732 / 101K=1,694 tok/s — ~8% decay over 14x context (KDA) ·
decode: c1 74 / c2 113 / c4 186 / c8 223 tok/s aggregate (NEXTN MTP-5, accept ~6.3/6). run-to-run ±2%. ·
KV pool 4,391,936 tok
@1M ctx · vision OK. Known gap: during a large prefill the first
concurrent decode drops to ~8% for the prefill window (~40s at 70K) — scheduler tuning TBD.
## Credits
-
@0xSero — sm120/sm121 patch bundle (the NoPE zero-rope FlashMLA path is the key unlock)
- SGLang team — day-0 glm-5.3-flash image
- zai-org — GLM-5.3-Flash open weights (MIT)