then...I ended up running tensor parallel, but deeepseek v4 flash has only one KV head, so the trick is to copy it to both machines, and split all the other layets.
The result: the best prefill of any setup (1,810 tok/s, +23% over one Mac) but the worst decode (32 tok/s). Every token needs 60+ syncs between the Macs.