i don't care how you get there, but the day you secure 2x DGX Spark is the endgame of local ai. you feel frontier intelligence right in front of your desk.
if a single RTX 3090 shows you the way, 2x DGX Spark will enlighten you. secure the boxes and you'll rarely miss closed models like claude or chatgpt.
it's been about 3 months with my 2x DGX Spark and i'm not making assumptions, i've run Qwen 3.8 Flash Next, DeepSeek V4 Flash and now GLM 5.3 Flash on them, and i'm still sitting here surprised by what these two nodes can do.
you can do some crazy shit with these boxes and they sip power, the two GPUs pull about 73W together while GLM decodes.
i'm really happy with these two boxes, and i hope i can add a 3rd one if nvidia shows grace.
loading GLM-5.3 Flash on my 2x DGX Spark, 256GB of unified memory across two boxes, NVIDIA's NVFP4 weights split over the ConnectX cable on stock vLLM
numbers so far:
> 320B total, 18B active, 204GB weights
> decode, MTP off: 15 tok/s, flat to 128K
> decode, MTP on (k=3): 19.5-31.7 tok/s
> code + math, MTP on: 26.7-31.7 tok/s
> 120K context, MTP on: 30.4 tok/s
> prefill, MTP off: 1,350-1,580 tok/s
> 128K prompt: first token in 84s
> thinking on, temp 1.0, live server
working on improving it now, testing quality and all, i've been waiting to try this model, moved home, got my wifi just yesterday and today it's loaded.