🧠⚡ Running GLM-5.2 at home the FULL 744B, all 256 experts, UNPRUNED across 4× NVIDIA DGX Spark (GB10).
200K context · MTP spec-decode · fp8 sparse-MLA KV · vLLM native multi-node
~30 tok/s single-stream, 60 tok/s @ 6 concurrent.
Full replicable recipe, open-sourced 👇
github.com/tonyd2wild/GLM-5.…