Running a precision machining business. UV laser processing · CNC machining · local AI compute Personal notes. Not a company account.

Seoul
I run precision machining operations. Building toward an AI-native manufacturing company. head full of: → agents watching production lines → LLMs reading quality reports → models learning from every defect → workflows for dark factories Personal notes. Not a company account.
1
3
1,314
Huge thanks to @MiaAI_lab. Your recipe gave my 8× DGX Spark setup a massive performance boost. Further tuning, SGLang’s bounded replay, and rhys’s SG18/RoCEnante stack brought it to the speeds below. Nearly two weeks with DeepSeek V4.1 Flash, and I’m spending more time putting it to work than tweaking it. Thanks to everyone sharing their work! The M5 Ultra looks seriously impressive too. Seems like there’s still plenty of room for optimization.
Great numbers on the new M5 Ultra. Here's Qwen3.8-Flash on 2x DGX Sparks: Prose on 4 streams: 146.0 tok/s Single stream prose - 56.8 tok/s Prefill is 3500 tok/s So basically the same?
3
3
16
1,898
DeepSeek-V4.1-Flash on 8x DGX Spark, SGLang TP8/EP8. 48 tok/s single-stream prose 179 tok/s aggregate at c8 ~4.4K tok/s prefill at 32K 2.6M-token KV pool, 256K context Official MXFP4 expert / FP8 dense weights. Vision + tools on. Huge credit to @MiaAI_lab. Their 3x Spark recipe took my single-stream decode from 38 to 48 tok/s. Two additional flags unlocked the concurrency and prefill gains. Details below.
Replying to @MiaAI_lab
You'll get even better performance if you have 4. Same repo for 3 and 4 sparks: github.com/MiaAI-Lab/DeepSee…
4
3
42
4,736
Starting from @MiaAI_lab's recipe (FlashInfer b12x for block-FP8 dense, NCCL buffer trim, fused MoE finalize off), I added two flags: 1. --enable-decoder-swa-bounded-replay Prefill +44–65%, decode −1–4%, outputs unchanged. It's documented in the V4.1 tech report, but I hadn't seen it used in a Spark recipe. 2. --min-free-slots-delay 1 SGLang classifies DSpark as DFlash-family and silently admits only 7 streams when max-running-requests >= 8. This enabled true 8-way admission and improved c8 aggregate by 43%. Final config: DSpark block 3, decode CUDA graphs bs 1–8, prefill graph off, chunked prefill 2048, KV fp8_e4m3 page 256, mem fraction 0.83, FlashInfer 0.6.18, official weights untouched. Tested and rejected on TP8: DSpark block 4/5, chunk 1024/4096/8192, NCCL 4/16 channels, cuDNN MXFP8, NCCL_IB_MERGE_NICS, FlashInfer 0.7, SPS, and ragged verify. The official DSpark block 5 default was 15% slower here. Single stream: prose 48 · JSON 62 · reasoning 63 · math 74 · coding 77 · counting 84 tok/s Thinking on: prose 61 tok/s At 128K context: 3.5K tok/s prefill, 43 tok/s decode Public 9-category prompt set, temperature 0, streaming, 3-run median.
203
Almost got my 40 tok/s GLM-5.3 743B on 8× DGX Spark: 38.8 tok/s single-stream prose 1,266 tok/s prefill @ 100K 110.7 tok/s aggregate @ c8 524K context The surprise: shorter speculation won. DFlash2 k7 → built-in MTP k2. For prose, the longer draft was costing more to verify than it returned. I was chasing 40. I'll take 38.8 for now :) Built on @Tech2Wild's Int4/Int8Mix quant and his 4× TP4 recipe as the starting point. Serving stack: eugr's spark-vllm-b12x image (local-inference-lab vLLM fork + b12x kernels).
Meanwhile on the big brother: GLM-5.3 743B (Int4-Int8Mix) on 8× DGX Spark. TP8. Still tuning, but already: c1 · temp 0 structured · 78 tok/s code · ~45 tok/s prose · 35 tok/s Prefill · ~1.1K tok/s 490K-token needle · PASS KV pool · 944K tokens Context · 512K DFlash2 k7 + NVFP4 KV on the b12x stack. Prose is still putting up a fight. More tuning tonight. Recipe drops when it stabilizes. Base lineage: @Tech2Wild's TP4 recipe Stack by @wrldsuksgo2mars · draft by @inco_ai
2
1
21
4,428
Serving details for the 8× DGX Spark GLM-5.3 run. HARDWARE / FABRIC 8× DGX Spark (GB10, 128 GB unified each), one head + seven workers MikroTik CRS804-4DDQ, 200G RoCE per node, MTU 9000 end-to-end (per @Tech2Wild's recipe gotcha) Each CX7 exposes two functions: enp1s0f0np0 / enP2p1s0f0np0 and rocep1s0f0 / roceP2p1s0f0. Separate /24 per function, both HCAs in NCCL_IB_HCA, RoCEv2 GID index 3 No PFC / ECN SOFTWARE eugr/spark-vllm-b12x:latest (2026-09-04): vLLM 0.1.dev20489 (local-inference-lab dev/jovian-judgement), b12x 1.3.0, PyTorch 2.13 / CUDA 13.0, NCCL 2.29.7 Driver 580.x, kernel 6.17-nvidia, 64 GB swap with vm.swappiness=100, min_free_kbytes=2 GB Model: Tech2wild/GLM-5.3-Int4-Int8Mix (743B, ~378 GB), local NVMe copy on every node SERVING CONFIG TP8 / PP1, one container per node, workers first, head last --quantization compressed-tensors, --attention-backend B12X --speculative-config {"method":"mtp","num_speculative_tokens":2,"attention_backend":"B12X"} --max-model-len 524288, --max-num-seqs 8, --max-num-batched-tokens 8192 --kv-cache-dtype fp8_ds_mla, --kv-cache-memory-bytes 36 GiB (pinned), --gpu-memory-utilization 0.84 --enable-prefix-caching, --async-scheduling, cudagraph_mode=FULL_AND_PIECEWISE Container: --memory 112g, --ipc host, --network host, /dev/infiniband, unlimited memlock THE TWO KNOBS THAT MATTERED MOST VLLM_ENABLE_ROCE_ALLREDUCE=1 — b12x RoCEnante one-shot RoCE all-reduce/all-gather for TP collectives instead of NCCL VLLM_B12X_MLA_CKV_GATHER=0 — removes the size cap on the MLA compressed-KV gather; long prefill was slower with the cap enabled WHERE THE 38.8 CAME FROM (same image, same benchmark protocol) v9 as-is, DFlash2 k=7: 19.5 tok/s prose speculation off: 20.0 built-in MTP k=2: 28.5 MTP k=2 + RoCE all-reduce: 38.1 / 38.8 / 38.0 across three sessions Why k=2 for prose: MTP k=2 accepts ~2.2–2.3 tokens/step at ~51 ms/step. Verification costs ~4.7 ms per drafted token here, so a 7-token draft spends more on verification than it returns on open-ended text. Code accepts more and can still favor longer drafts. RESULTS Single-stream prose: 38.1–38.8 tok/s median of 10 runs, 250 output tokens, wall-clock incl. TTFT Prefill: 1,406 / 1,324 / 1,266 tok/s at 32K / 70K / 100K Concurrency, aggregate tok/s: c1 → 36.6 c2 → 58.4 c4 → 62.4 c8 → 110.7 TTFT: 250–610 ms --max-num-seqs 8 vs 4: no measurable single-stream penalty Cold load ~13 min, warm ~8 min GOTCHAS Prefill decays with node uptime. The same launcher went from ~1,400 tok/s after reboot to ~600 later in the day. Restarting vLLM doesn't recover it; rebooting the nodes does. The bad state shows page compaction failing on all eight nodes (compact_fail ≈ compact_stall, no order ≥10 free blocks). Correlation only; cause not proven. Stock vLLM doesn't boot this model on GB10. The b12x sm12x kernels (sparse MLA + fp8 KV path) in the image above are required. The fabric is lossy: no PFC/ECN. It has been fine for pure inference and RoCE error counters stay flat. Training-style all-to-all does produce tx drops and rising packet_seq_err. METHOD Single stream: 10 unique prose prompts (6 Korean, 4 English), 250 output tokens each; median wall-clock throughput Prefill: unique 32K / 70K / 100K-token documents, single request Concurrency: N simultaneous 250-token prose streams; aggregate = sum of per-stream tok/s
3
3
393
Meanwhile on the big brother: GLM-5.3 743B (Int4-Int8Mix) on 8× DGX Spark. TP8. Still tuning, but already: c1 · temp 0 structured · 78 tok/s code · ~45 tok/s prose · 35 tok/s Prefill · ~1.1K tok/s 490K-token needle · PASS KV pool · 944K tokens Context · 512K DFlash2 k7 + NVFP4 KV on the b12x stack. Prose is still putting up a fight. More tuning tonight. Recipe drops when it stabilizes. Base lineage: @Tech2Wild's TP4 recipe Stack by @wrldsuksgo2mars · draft by @inco_ai
6
1
11
3,956
Full serving setup for the 743B run: · Model: QuantTrio/GLM-5.3-Int4-Int8Mix (377GB, compressed-tensors) · Draft: incoai/GLM-5.3-DFlash2, k=7 · Engine: b12x-lineage vLLM (sm121), TP8 over 8 nodes, dual-rail RoCE, NCCL 8 channels · KV cache: nvfp4_ds_mla + fp8-skip on sliding-window layers, pinned --kv-cache-memory at 36GiB/node → ~944K-token pool · ctx 524,288 · chunked prefill 8192 · gpu-mem-util 0.86 Notes from this setup (8× GB10, unified memory): I pinned KV for reproducibility. With auto GMU, the resulting pool size varied between boots in my setup. In my tests, long-context prefill performance dropped as reported free memory per node fell below ~8GiB, with stalls appearing below ~5GiB. I ladder-tested KV pins from 26→44GiB; 36GiB gave the best balance in those runs. I could reach a ~1.18M-token pool with a 44GiB pin + smaller prefill chunks, but measured ~28% lower prefill throughput in that configuration. Not worth the tradeoff for my use. Draft TP 1 vs 8: I couldn't measure a meaningful difference in these runs. On this checkpoint, at temp 1, MTP acceptance fell to ~3%. DFlash2 maintained substantially higher acceptance under the same test. Prose is currently sitting around 30% draft acceptance in my tests, so that's where I'm focusing next.
2
205
You have 14 seconds to guess. No cheating :) How many tok/s does this look like? Hint: prose is the slow lane for speculative decoding. Structured workloads push 240+. GLM-5.3-Flash · uncensored 2× RTX PRO 6000 360K context · fully local vLLM b12x · DFlash2 k7 · NVFP4 KV · EXL3 4bpw Quant by @softpoo (from @OrcaRouter's FP8) Stack & recipe by @wrldsuksgo2mars Draft model by @inco_ai h/t @TechMDAI 🙏
11
1
37
10,548
TL;DR: 230 tok/s structured · 350 tok/s aggregate code @ c4 · 360K ctx · uncensored · fully local Performance (temp 0, measured via usage tokens) Single stream (c1): · structured: 230 tok/s · code: 154 tok/s 4 concurrent users (aggregate): · structured: 589 tok/s · code: 350 tok/s · prose: 236 tok/s c8 peaks at 595 tok/s aggregate. Prefill: 2.5–2.8K tok/s. 332K-token needle test: PASS. Full serving setup: · Model: neko-legends/GLM-5.3-Flash-Uncensored-EXL3, 4bpw, mcg codebook · Draft: incoai/GLM-5.3-Flash-DFlash2, k=7 · Engine: @wrldsuksgo2mars's b12x vLLM fork, TP2 + DCP2 (decode context parallel) · KV cache: nvfp4_ds_mla with pinned --kv-cache-memory → 957K-token KV pool · ctx 368,640 · chunked prefill 1024 · 98.3% of 2×96GB VRAM in use Two notes from this setup: mul1-codebook EXL3 quants didn't run on this path; mcg did. Prefill chunks >1024 stalled at long context in my setup; 1024 was stable. And the answer to the video? 97 tok/s. Prose typically lands around 85–105 tok/s depending on the prompt. How close were you? :)
5
752
Hopefully this little duck can become the foreman of my factory someday. First lesson: how not to get run over by the AMRs. 🦆
We built a small biped robot you can teach new tricks to. Train it in simulation, run it on the real thing. Meet Microduck 🦆 $399, shipping before Christmas. pollen-robotics.com/microduc… github.com/pollen-robotics/m…
2
271
GLM-5.3-Flash on 8× DGX Spark (GB10). Official FP8 · TP8 · 1M ctx · SGLang ⚡ 1,842 tok/s prefill 🚀 74 tok/s single-stream decode 🚀 223 tok/s aggregate @ c8 🧠 4.39M-token KV pool The number that matters is prefill decay. Cache-busted: 7K → 1,842 tok/s 37K → 1,732 tok/s 101K → 1,694 tok/s 14× context, ~8% drop. 34/45 layers are linear attention. Long-context prefill barely degrades. Decode (NEXTN MTP-5): c1 74 · c2 113 · c4 186 · c8 223 tok/s agg Vision working. Measured to 101K; serving 1M. SM121 notes: vLLM NoPE/FP8 KV crash, SGLang cutlass needed Triton, RDMA needed IPC_LOCK, 2048 chunks beat 8192. Built on @0xSero's SM121 bundle + @sgl_project day-0 image + @Zai_org weights. Full repro below 👇
7
4
30
2,767
# GLM-5.3-Flash (FP8) on 8x DGX Spark — TP8 serving info (reproducible) ## TL;DR World-first (AFAIK) GLM-5.3-Flash on 8x DGX Spark (GB10/sm_121a), TP8, **official zai-org FP8 checkpoint** (not NVFP4), 1M context, 4.39M-token KV pool, NEXTN MTP-5, vision working. Built on @0xSero's sm121 bundle — extended from TP4/NVFP4 to TP8/FP8 with two extra workarounds. ## Hardware - 8x NVIDIA DGX Spark (GB10, sm_121a, 121.69 GiB unified memory each) - RoCEv2 dual-rail mesh: 2x ConnectX-7 per node (200GbE effective/node), MikroTik CRS804 - 1 GPU per node → TP_SIZE = EP_SIZE = NNODES = 8 ## Stack - Base image: lmsysorg/sglang:glm-5.3-flash (day-0 GLM-5.3 image) - Patches: github.com/0xSero/glm-5.3-fl… (6 baked patches — glm5_next model, NEXTN draft, sm120 quant utils, modelopt quant, flash_mla_sm120 w/ NoPE zero-rope path, dsa_backend) - Rebuilt with TORCH_CUDA_ARCH_LIST=12.1a / CUTE_DSL_ARCH=sm_121a, linux/arm64 - Model: zai-org/GLM-5.3-Flash (native FP8, 306 GiB, 62 shards) — every node holds a full local copy ## Why vLLM does NOT work (as of 2026-08-27) GLM-5.3-Flash MLA is NoPE (qk_rope_head_dim=0). The only attention backend vLLM offers on sm_121 (FLASHINFER_MLA_SPARSE_SM120) force-selects the fp8_ds_mla KV format, whose kernel hard-codes pe_dim=64 → `concat_and_cache_mla: pe_dim must be 64 for fp8_ds_mla`, crash at ~55% load, 3/3 repro. Not bypassable by --kv-cache-dtype (auto is overridden by the backend). ## Two extra workarounds needed for FP8 + TP8 1. **SGLang Fp8MoEMethod bug**: with `--moe-runner-backend flashinfer_cutlass` (the bundle default), create_moe_runner() silently skips runner init for cutlass ("else: pass # TODO") → `AttributeError: 'Fp8MoEMethod' object has no attribute 'runner'`. Workaround: `--moe-runner-backend triton`. 2. **RDMA memory registration**: add `--cap-add IPC_LOCK --ulimit memlock=-1:-1` to docker run, or NCCL dies with `ibv_reg_mr_iova2 failed: Cannot allocate memory`. Also: `--enable-prefill-cp` (zigzag) crashes in cuda-graph capture (deep-ep invalid resource handle) on this combo — do not use. ## Launch (per node, rank 0..7, workers use same cmd) docker run -d --name glm53 --gpus all --network host --ipc host --shm-size 32gb \ --cap-add IPC_LOCK --ulimit memlock=-1:-1 --device /dev/infiniband:/dev/infiniband \ -v /var/tmp/models:/cache/huggingface \ -e NCCL_NET=IB -e NCCL_IB_HCA=<your_rocev2_hcas> -e NCCL_SOCKET_IFNAME=<your_ifaces> \ -e NCCL_IB_GID_INDEX=3 -e NCCL_MAX_NCHANNELS=8 -e NCCL_MIN_NCHANNELS=8 -e NCCL_BUFFSIZE=16777216 \ <sm121-patched-image> \ python3 -m sglang.launch_server \ --model-path /cache/huggingface/GLM-5.3-Flash-FP8 \ --served-model-name glm-5.3-flash --host 0.0.0.0 --port 8210 \ --tp-size 8 --ep-size 8 --nnodes 8 --node-rank $RANK \ --dist-init-addr <head_mesh_ip>:27000 \ --context-length 1048576 \ --attention-backend dsa \ --dsa-prefill-backend flashinfer_sparse_mla --dsa-decode-backend flashinfer_sparse_mla \ --linear-attn-backend triton \ --kv-cache-dtype fp8_e4m3 \ --moe-runner-backend triton \ --disable-shared-experts-fusion \ --chunked-prefill-size 2048 --max-prefill-tokens 2048 \ --max-running-requests 8 --mem-fraction-static 0.83 \ --cuda-graph-max-bs-decode 8 \ --speculative-algorithm NEXTN --speculative-num-steps 5 \ --speculative-eagle-topk 1 --speculative-num-draft-tokens 6 --speculative-adaptive \ --trust-remote-code (note: single value only for --served-model-name; NCCL ifaces are cluster-specific) ## Non-obvious tuning finds - **--chunked-prefill-size 2048 doubled long prefill** vs the "bigger is better" intuition: 70K-token prefill = 887 tok/s @8192 → **1,748 tok/s @2048** (1024 regresses to 1,575). Bonus: prefill window halves, so concurrent-decode starvation window halves too. - mem-fraction-static 0.83, not 0.90: GB10 has a boot-time usable ceiling (~101.6/121.69 GiB). - NCCL_MAX/MIN_NCHANNELS=8 (from our GLM-5.2 tuning — 2.6x prefill vs 4 channels on this mesh). - cold first request ~232 tok/s (triton JIT) — warm up before judging numbers. - Engine startup ~6 min (weights 235s + cuda-graph capture). ## Numbers (see screenshot) prefill (cache-busted): 7K=1,842 / 37K=1,732 / 101K=1,694 tok/s — ~8% decay over 14x context (KDA) · decode: c1 74 / c2 113 / c4 186 / c8 223 tok/s aggregate (NEXTN MTP-5, accept ~6.3/6). run-to-run ±2%. · KV pool 4,391,936 tok @1M ctx · vision OK. Known gap: during a large prefill the first concurrent decode drops to ~8% for the prefill window (~40s at 70K) — scheduler tuning TBD. ## Credits - @0xSero — sm120/sm121 patch bundle (the NoPE zero-rope FlashMLA path is the key unlock) - SGLang team — day-0 glm-5.3-flash image - zai-org — GLM-5.3-Flash open weights (MIT)
1
4
569
▎ Correction to the launch command: add ▎ --reasoning-parser glm45 --tool-call-parser glm47 ▎ — missing from my original paste. Without these, tool calls leak as raw <tool_call> XML in content and agentic clients silently break. No effect on the benchmark numbers — parsers are post-processing on the output text, not in the inference path.
1
134
# GLM-5.3-Flash (FP8) on 8x DGX Spark — TP8 serving info (reproducible) ## TL;DR World-first (AFAIK) GLM-5.3-Flash on 8x DGX Spark (GB10/sm_121a), TP8, **official zai-org FP8 checkpoint** (not NVFP4), 1M context, 4.39M-token KV pool, NEXTN MTP-5, vision working. Built on @0xSero's sm121 bundle — extended from TP4/NVFP4 to TP8/FP8 with two extra workarounds. ## Hardware - 8x NVIDIA DGX Spark (GB10, sm_121a, 121.69 GiB unified memory each) - RoCEv2 dual-rail mesh: 2x ConnectX-7 per node (200GbE effective/node), MikroTik CRS804 - 1 GPU per node → TP_SIZE = EP_SIZE = NNODES = 8 ## Stack - Base image: lmsysorg/sglang:glm-5.3-flash (day-0 GLM-5.3 image) - Patches: github.com/0xSero/glm-5.3-fl… (6 baked patches — glm5_next model, NEXTN draft, sm120 quant utils, modelopt quant, flash_mla_sm120 w/ NoPE zero-rope path, dsa_backend) - Rebuilt with TORCH_CUDA_ARCH_LIST=12.1a / CUTE_DSL_ARCH=sm_121a, linux/arm64 - Model: zai-org/GLM-5.3-Flash (native FP8, 306 GiB, 62 shards) — every node holds a full local copy ## Why vLLM does NOT work (as of 2026-08-27) GLM-5.3-Flash MLA is NoPE (qk_rope_head_dim=0). The only attention backend vLLM offers on sm_121 (FLASHINFER_MLA_SPARSE_SM120) force-selects the fp8_ds_mla KV format, whose kernel hard-codes pe_dim=64 → `concat_and_cache_mla: pe_dim must be 64 for fp8_ds_mla`, crash at ~55% load, 3/3 repro. Not bypassable by --kv-cache-dtype (auto is overridden by the backend). ## Two extra workarounds needed for FP8 + TP8 1. **SGLang Fp8MoEMethod bug**: with `--moe-runner-backend flashinfer_cutlass` (the bundle default), create_moe_runner() silently skips runner init for cutlass ("else: pass # TODO") → `AttributeError: 'Fp8MoEMethod' object has no attribute 'runner'`. Workaround: `--moe-runner-backend triton`. 2. **RDMA memory registration**: add `--cap-add IPC_LOCK --ulimit memlock=-1:-1` to docker run, or NCCL dies with `ibv_reg_mr_iova2 failed: Cannot allocate memory`. Also: `--enable-prefill-cp` (zigzag) crashes in cuda-graph capture (deep-ep invalid resource handle) on this combo — do not use. ## Launch (per node, rank 0..7, workers use same cmd) docker run -d --name glm53 --gpus all --network host --ipc host --shm-size 32gb \ --cap-add IPC_LOCK --ulimit memlock=-1:-1 --device /dev/infiniband:/dev/infiniband \ -v /var/tmp/models:/cache/huggingface \ -e NCCL_NET=IB -e NCCL_IB_HCA=<your_rocev2_hcas> -e NCCL_SOCKET_IFNAME=<your_ifaces> \ -e NCCL_IB_GID_INDEX=3 -e NCCL_MAX_NCHANNELS=8 -e NCCL_MIN_NCHANNELS=8 -e NCCL_BUFFSIZE=16777216 \ <sm121-patched-image> \ python3 -m sglang.launch_server \ --model-path /cache/huggingface/GLM-5.3-Flash-FP8 \ --served-model-name glm-5.3-flash --host 0.0.0.0 --port 8210 \ --tp-size 8 --ep-size 8 --nnodes 8 --node-rank $RANK \ --dist-init-addr <head_mesh_ip>:27000 \ --context-length 1048576 \ --attention-backend dsa \ --dsa-prefill-backend flashinfer_sparse_mla --dsa-decode-backend flashinfer_sparse_mla \ --linear-attn-backend triton \ --kv-cache-dtype fp8_e4m3 \ --moe-runner-backend triton \ --disable-shared-experts-fusion \ --chunked-prefill-size 2048 --max-prefill-tokens 2048 \ --max-running-requests 8 --mem-fraction-static 0.83 \ --cuda-graph-max-bs-decode 8 \ --speculative-algorithm NEXTN --speculative-num-steps 5 \ --speculative-eagle-topk 1 --speculative-num-draft-tokens 6 --speculative-adaptive \ --trust-remote-code (note: single value only for --served-model-name; NCCL ifaces are cluster-specific) ## Non-obvious tuning finds - **--chunked-prefill-size 2048 doubled long prefill** vs the "bigger is better" intuition: 70K-token prefill = 887 tok/s @8192 → **1,748 tok/s @2048** (1024 regresses to 1,575). Bonus: prefill window halves, so concurrent-decode starvation window halves too. - mem-fraction-static 0.83, not 0.90: GB10 has a boot-time usable ceiling (~101.6/121.69 GiB). - NCCL_MAX/MIN_NCHANNELS=8 (from our GLM-5.2 tuning — 2.6x prefill vs 4 channels on this mesh). - cold first request ~232 tok/s (triton JIT) — warm up before judging numbers. - Engine startup ~6 min (weights 235s + cuda-graph capture). ## Numbers (see screenshot) prefill (cache-busted): 7K=1,842 / 37K=1,732 / 101K=1,694 tok/s — ~8% decay over 14x context (KDA) · decode: c1 74 / c2 113 / c4 186 / c8 223 tok/s aggregate (NEXTN MTP-5, accept ~6.3/6). run-to-run ±2%. · KV pool 4,391,936 tok @1M ctx · vision OK. Known gap: during a large prefill the first concurrent decode drops to ~8% for the prefill window (~40s at 70K) — scheduler tuning TBD. ## Credits - @0xSero — sm120/sm121 patch bundle (the NoPE zero-rope FlashMLA path is the key unlock) - SGLang team — day-0 glm-5.3-flash image - zai-org — GLM-5.3-Flash open weights (MIT)
218
The new M5 Ultra Mac Studio looks like an absolute beast for local AI. 512GB unified memory. 1.2TB/s memory bandwidth. Neural Accelerators in every GPU core. As someone running an 8× DGX Spark cluster, this immediately got my attention. There's really only one thing I want to see now: Long-context prefill. Decode was never what held me back on Apple silicon. Prefill was. If Apple has significantly closed that gap with the new Neural Accelerators, this thing is going to be a monster. 512GB. 1.2TB/s. One box. Now show me the benchmarks.
227
My workflow is now almost entirely local. DeepSeek-V4-Flash on 2× RTX PRO 6000. GLM-5.2 on the 8× DGX Spark cluster. The biggest surprise wasn't the models. It was crossing ~244 tok/s. Once latency stops interrupting your train of thought, everything changes. You stop waiting. You just keep thinking.
DeepSeek-V4-Flash at 243 tok/s on 2× RTX PRO 6000. Benchmarks are nice. Real work is better. As a test, I had it build three terminal games in Claude Code. One crashed on launch, so I pasted the traceback back in and let it debug its own code. Watch it reason through a classic curses corner case and fix it. Unedited. 2:21. 0:00 - Traceback 0:20 - Reasoning 1:30 - Fix All local. No cloud.
4
4
49
6,291
403 tok/s on an 8× DGX Spark cluster. DeepSeek-V4-Flash-0731 GB10 · 200G RoCE · TP=8 Measured (usage-metered): • 403 tok/s aggregate @ c32 • 88 tok/s single stream • 1M context • 4.17M-token KV pool • max_num_seqs 64 • Zero failures The interesting part wasn't the throughput. DSpark speculative decoding inverted under concurrency. At c1 it delivered a 2.4× speedup over the no-spec baseline. By c32 it was ~9% slower. Once the batch fills, the acceptance economics flip. My 2× RTX PRO 6000 workstation still wins single-stream latency (244 vs 88 tok/s). The Spark cluster wins throughput. Different tools for different workloads.
4
1
35
2,912
Full recipe for the 8× GB10 setup: BUILD vLLM PR #41834 (SM121 sparse MLA + DSpark for GB10). Built from source on one node with: • TORCH_CUDA_ARCH_LIST=12.1a • FlashInfer 0.6.15.post1 (pinned) Then docker save/load to the other seven nodes. Build time: ~1 hour on a GB10. --- TOPOLOGY Native vLLM multi-node (no Ray). Each node runs: vllm serve --nnodes 8 --node-rank R --master-addr <head> Ranks >0 use --headless. TP=8 across the nodes over 200G RoCE. --- KEY SETTINGS • FP8 KV cache • max-model-len 1048576 • max-num-seqs 64 • Prefix caching disabled • --speculative-config '{"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}' num_speculative_tokens must be 5 (the DSpark block size). --- THINGS THAT COST US TIME • GMU-based sizing didn't work reliably here. We reserve KV memory explicitly (40 GiB per node) instead. • NCCL_IB_PCI_RELAXED_ORDERING=0 is required. Otherwise RoCE silently falls back to TCP. • Raise ulimit -n for 8-node multiprocessing. • The max_num_seqs > 4 crash from my RTX PRO 6000 recipe does NOT exist in this branch. max_num_seqs=64 ran clean. --- Validation followed the same sanity ladder: Health → Deterministic text → Usage-metered throughput → Needle (128K / 384K) → Tool calling
1
175