God-fearing husband & father. AI Engineer. Fine-tuning + day-zero EXL3 & GGUF quants + benchmarks • Local LLMs • honest tok/s • fitting big models on small GPUs

Pinned Tweet
LADIES AND GENTLEMEN, after 3 days of blessings, a 2.5 hour drive and a total of $4,500 in donations from this amazing blessing of a community I HAVE ACQUIRED THE DGX SPARK!! 😩😁😌🚀🚀🚀 Thank you to everyone who contributed and shared my post, I’m forever grateful for every single one of you!! I will be pumping out so much more quants and recipes for yall so stay tuned!! The crazy part about this is I got it for $4650, when it was posted at $4800 from Facebook marketplace, so I only needed to add $150 of my own cash!! 🤯😩 GOD IS SO GOOD!! A very special thank you to the most high, Jesus Christ for the provision and blessings brought my way. I’m forever grateful and will use this to continue glorifying your kingdom! 🙏
A lot of people in this community were genuinely shocked when they learned this: I do not own a DGX Spark. Much of the Spark-specific work I’ve released, including day-zero EXL3 and GGUF quants, vllm-exl3 development, custom kernels, SixCat evaluations, and reproducible serving recipes, has only been possible because incredible people in this community have loaned me access to their own hardware and clusters. I’m deeply grateful for that support. But borrowed access is temporary and unpredictable. I often have a new model, fix, or benchmark ready, then have to wait until someone else’s machine is available before I can validate and publish it properly. So I’m raising **$5,000 toward a DGX Spark of my own.** 🙏 𝗪𝗛𝗔𝗧 𝗔 𝗗𝗘𝗗𝗜𝗖𝗔𝗧𝗘𝗗 𝗦𝗣𝗔𝗥𝗞 𝗪𝗢𝗨𝗟𝗗 𝗘𝗡𝗔𝗕𝗟𝗘 → More true day-zero quantized model releases → Immediate hardware validation instead of projected results → Continuous vllm-exl3 and ExLlamaV3 kernel development → SixCat quality, speed, concurrency, and long-context benchmarks → Reproducible one-Spark recipes the community can run themselves → Faster debugging when new models or inference engines break It would turn a lot of: “I built it, but I need someone with a Spark to test it” into: “I built it, tested it, documented it, and released the recipe.” 𝗠𝗬 𝗣𝗨𝗕𝗟𝗜𝗖 𝗪𝗢𝗥𝗞 𝗦𝗧𝗔𝗬𝗦 𝗣𝗨𝗕𝗟𝗜𝗖 This is not an access fee, preorder, or paywall. The models, recipes, benchmarks, fixes, and tools will remain public. This is completely optional support for the hardware behind that work. I’m planning to place the order within 48 hours and will personally cover whatever the fundraiser does not. For complete transparency, I’ll publish: → The purchase receipt with private details removed → Photos of the actual DGX Spark → Total community contributions → Platform and payment fees → Final purchase cost → The amount I contributed myself To everyone who has loaned me hardware, funded compute, tested a build, submitted a PR, shared my work, or simply encouraged me: thank you. I genuinely could not have pushed this work as far as I have without this community. A contribution would mean a lot. A repost helps just as much. ❤️ ko-fi.com/victorangelcruz/go…
66
11
434
26,046
🎶 For every time you told me no For every time you told me yes I give thanks Lo and behold You had a plan Like Gideon and three hundred men I will be strong Speak life I'm bold I abide In the Most High Want everybody to know Him The one who's keeping me going Who wakes me up in the morning I want everybody to know He turned my mourning to dancing He's in my coming and going The difference from thinking and knowing So I will not worry what today brings or tomorrow I won't let time put a question on what I know All that shines and glitters it is not gold But I'm longing for your gifts bring no sorrow All that shines ain't glitter 🎵 No Sorrow - @SSTEDI I thank you lord for all that you have provided in times of need. I thank you holy lord of lords for the provision and the opportunity to glorify your name in all I dedicate myself to. Thank you for those you’ve put in my way who have been able to bless me in so many ways, look over them and give to them the same provision you’ve provided to me 100 fold. In Jesus name I pray, amen.
2
1
7
500
Im just blown away, 39 donations, $2,965 on Ko-fi. Plus the $1,500 BTC = $4,465 total which is 89.30% of the $5,000 goal. ONLY $535 left !! Thank you all so much! Now the fun part: Anyone know where I can find a DGX Spark or willing to sell me one for regular price and not an arm and a leg?! 😅 Even though im hunting everywhere for one it seems I cant find any at this price range so im going to be patient! But I will keep you all posted! 🫡 Feel free to keep reposting for me we are almost there!! 🚀 ko-fi.com/victorangelcruz/go… - Ko-fi: $2,965 — 59.30% of goal - BTC: $1,500 — 30.00% of goal BTC Address: 3GT8KQuh9xagVbFsGjHSvaNAUW5bXsty3i
Even with the DGX Spark sold out everywhere I can’t believe we are already at $3,674 total! 73.48% of the $5,000 goal! With only $1,326 left (26.52%)!! I’m extremely grateful to you all! I’m also keeping my eyes peeled on MicroCenter stock (Columbia, OH got 6 more earlier today) and on eBay!! Tomorrow is my final day of the fund! If anyone is willing to contribute or share it’ll be greatly appreciated! Thank you!! - Ko-fi: $2,174 — 43.48% of goal, 59.17% of everything raised ko-fi.com/victorangelcruz/go… - BTC: $1,500 — 30.00% of goal, 40.83% of everything raised 3GT8KQuh9xagVbFsGjHSvaNAUW5bXsty3i
4
1
30
1,834
RTX PRO 6000 96GB + DGX Spark owners rejoice! 🔥 You can now run the highest quality Xiaomi’s MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process 🚀 309B total parameters. Only ~15B active per token. Xiaomi’s published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5. And on ONE RTX PRO 6000: 184.1 tok/s p50 with DFlash 49.6 tok/s without drafting 2,258+ tok/s prefill 321.9 tok/s confirmed aggregate @ C=8 𝗣𝗜𝗖𝗞 𝗬𝗢𝗨𝗥 𝗖𝗔𝗥𝗗 RTX PRO 6000 96GB: 2.20 bpw SAGE-EXL3 86.94 GB DGX Spark 128GB: 2.50 bpw SAGE-EXL3 98.48 GB The recipe includes separate profiles for both systems. The 184.1 tok/s result is specifically from the RTX PRO 6000; Spark DFlash gains are currently lower and more prompt-dependent. 𝗙𝗜𝗗𝗘𝗟𝗜𝗧𝗬 𝗥𝗘𝗖𝗘𝗜𝗣𝗧𝗦 Across two completely untouched holdout sets: 2.20 bpw: 82.16–82.64% top-1 agreement 0.1926–0.1986 mean KLD 2.50 bpw: 83.76–87.30% top-1 agreement 0.1086–0.2055 mean KLD Top-1 agreement means the EXL3 pack selected the SAME most-likely next token as Xiaomi’s reference. KLD compares the entire next-token probability distribution. Lower means the quant tracks the reference more closely. 𝗔𝗡 𝗜𝗠𝗣𝗢𝗥𝗧𝗔𝗡𝗧 𝗗𝗘𝗧𝗔𝗜𝗟 Xiaomi does not publish a full BF16 expert checkpoint. The 303B routed-expert bank already ships in MXFP4. Attention ships in block-FP8, with embeddings, norms and the output head in BF16. So these fidelity numbers compare against Xiaomi’s actual released mixed-precision checkpoint running through its official implementation, not against a hidden full-BF16 teacher. This is effectively EXL3 agreement with the best public source that exists. 𝗛𝗢𝗪 𝗖𝗟𝗢𝗦𝗘 𝗜𝗦 𝗙𝗟𝗔𝗦𝗛 𝗧𝗢 𝗧𝗛𝗘 𝗙𝗥𝗢𝗡𝗧𝗜𝗘𝗥? Xiaomi’s published numbers: AutomationBench: MiMo Flash 52.3 Claude Opus 5 50.3 GPT-5.6 Sol 45.8 Terminal-Bench 2.1: MiMo Flash 87.6 Claude Opus 5 89.1 GPT-5.6 Sol 88.8 OSWorld: MiMo Flash 80.8 Claude Opus 5 83.4 GPT-5.6 Sol 83.0 VisualCoding: MiMo Flash 71.5 Claude Opus 5 70.0 GPT-5.6 Sol 73.4 Artificial Analysis has not scored Flash independently yet. The larger MiMo-V2.6-Pro sibling scores 46 on the AA Intelligence Index, tied with Grok 4.7 and sitting beside GLM-5.3 at 45. This release is the text backbone. The original DFlash speculative drafter is repaired and wired through the recipe; vision and audio towers are not included yet. Huge credit to @XiaomiMiMo team for the model and to @turboderp_ / ExLlamaV3 for the EXL3 engine and format. @XiaomiMiMoDevs EXL3 weights + fidelity results: huggingface.co/vcruz305/MiMo… Complete serving recipe: github.com/vcruz305/MiMo-V2.…
4
10
70
5,159
Amazing work by Chris getting my TP4 (4x Sparks) Deepseekv4.1-Flash-EXL3 that’s 4.75bpw to an insane 64.8 Tok/s!!! 🤯 @cfontes made many improvements to my Vllm-exl3 plugin for which I’m extremely grateful for! Thank you so much! 🙏 Keep in mind @deepseek_ai’s Deepseek v4.1-Flash scored higher than the most recently released GPT-6 Luna on max, Sonnet 5 on max and 1 point less than the new GPT-6 Sol on medium!! On 4 DGX Sparks! These frontier labs better catch up or they will get left behind by local compute maxxies!! 🚀
Replying to @deepseek
@DeepSeek-V4.1-Flash (EXL3 4.75bpw) on 4× DGX Sparks: big update since our last champion. 📈 Results • Single-stream decode: 28.2 → 36 tok/s • Code, 1 stream: 64.8 tok/s • Code, 10 concurrent: 166.7 tok/s aggregate YES. 166.7 tok/s on code (aggregate)! 🔧 What changed 1. Tensor-parallel MoE instead of expert-parallel. With EP, each token's experts land unevenly across nodes, so every layer waited on the busiest one: 200–500µs stalls after each MoE all-reduce. EXL3's Hadamard transforms work in 128-wide blocks, so we split the 2304 intermediate dim as 640/640/512/512. Step time 63 → 56ms. 2. Padded MoE loops bounded by live rows. The kernel walked all 32 padded rows even when 3 were live. Now sizing for c10 costs c1 nothing. Output is bit-identical. 3. Grouped-expert MoE kernel (the big one). The expert kernels ran m16n8k16 tensor-core MMAs but filled only 1 of 16 A-rows. Rows that share an expert now go into those idle rows, so each weight tile is decoded once per expert instead of once per row. Output is bit-identical. c10 aggregate +58%. 4. Speculative length scales with load. DSpark was trained for 5-token blocks, and on code it accepts ~4 tokens per step, but every verified row costs step time. So: k=4 for 1–4 concurrent requests, k=3 above. It lands within ~3% of the best fixed k at every concurrency. 🔗 Open PRs to vllm-exl3 (thanks @ViC305): #36 TP-MoE aligned split: github.com/vcruz305/vllm-exl… #37 Padded-MoE live-row bounds: github.com/vcruz305/vllm-exl… #39 Grouped-expert MoE kernel: github.com/vcruz305/vllm-exl… Recipe with configs + full benchmark tables: github.com/vcruz305/DeepSeek… Next up: a one-shot RDMA all-reduce, aiming for ~40 tok/s prose at c1.
6
20
1,739
Even with the DGX Spark sold out everywhere I can’t believe we are already at $3,674 total! 73.48% of the $5,000 goal! With only $1,326 left (26.52%)!! I’m extremely grateful to you all! I’m also keeping my eyes peeled on MicroCenter stock (Columbia, OH got 6 more earlier today) and on eBay!! Tomorrow is my final day of the fund! If anyone is willing to contribute or share it’ll be greatly appreciated! Thank you!! - Ko-fi: $2,174 — 43.48% of goal, 59.17% of everything raised ko-fi.com/victorangelcruz/go… - BTC: $1,500 — 30.00% of goal, 40.83% of everything raised 3GT8KQuh9xagVbFsGjHSvaNAUW5bXsty3i
A lot of people in this community were genuinely shocked when they learned this: I do not own a DGX Spark. Much of the Spark-specific work I’ve released, including day-zero EXL3 and GGUF quants, vllm-exl3 development, custom kernels, SixCat evaluations, and reproducible serving recipes, has only been possible because incredible people in this community have loaned me access to their own hardware and clusters. I’m deeply grateful for that support. But borrowed access is temporary and unpredictable. I often have a new model, fix, or benchmark ready, then have to wait until someone else’s machine is available before I can validate and publish it properly. So I’m raising **$5,000 toward a DGX Spark of my own.** 🙏 𝗪𝗛𝗔𝗧 𝗔 𝗗𝗘𝗗𝗜𝗖𝗔𝗧𝗘𝗗 𝗦𝗣𝗔𝗥𝗞 𝗪𝗢𝗨𝗟𝗗 𝗘𝗡𝗔𝗕𝗟𝗘 → More true day-zero quantized model releases → Immediate hardware validation instead of projected results → Continuous vllm-exl3 and ExLlamaV3 kernel development → SixCat quality, speed, concurrency, and long-context benchmarks → Reproducible one-Spark recipes the community can run themselves → Faster debugging when new models or inference engines break It would turn a lot of: “I built it, but I need someone with a Spark to test it” into: “I built it, tested it, documented it, and released the recipe.” 𝗠𝗬 𝗣𝗨𝗕𝗟𝗜𝗖 𝗪𝗢𝗥𝗞 𝗦𝗧𝗔𝗬𝗦 𝗣𝗨𝗕𝗟𝗜𝗖 This is not an access fee, preorder, or paywall. The models, recipes, benchmarks, fixes, and tools will remain public. This is completely optional support for the hardware behind that work. I’m planning to place the order within 48 hours and will personally cover whatever the fundraiser does not. For complete transparency, I’ll publish: → The purchase receipt with private details removed → Photos of the actual DGX Spark → Total community contributions → Platform and payment fees → Final purchase cost → The amount I contributed myself To everyone who has loaned me hardware, funded compute, tested a build, submitted a PR, shared my work, or simply encouraged me: thank you. I genuinely could not have pushed this work as far as I have without this community. A contribution would mean a lot. A repost helps just as much. ❤️ ko-fi.com/victorangelcruz/go…
2
3
35
4,930
Cruz retweeted
A lot of people in this community were genuinely shocked when they learned this: I do not own a DGX Spark. Much of the Spark-specific work I’ve released, including day-zero EXL3 and GGUF quants, vllm-exl3 development, custom kernels, SixCat evaluations, and reproducible serving recipes, has only been possible because incredible people in this community have loaned me access to their own hardware and clusters. I’m deeply grateful for that support. But borrowed access is temporary and unpredictable. I often have a new model, fix, or benchmark ready, then have to wait until someone else’s machine is available before I can validate and publish it properly. So I’m raising **$5,000 toward a DGX Spark of my own.** 🙏 𝗪𝗛𝗔𝗧 𝗔 𝗗𝗘𝗗𝗜𝗖𝗔𝗧𝗘𝗗 𝗦𝗣𝗔𝗥𝗞 𝗪𝗢𝗨𝗟𝗗 𝗘𝗡𝗔𝗕𝗟𝗘 → More true day-zero quantized model releases → Immediate hardware validation instead of projected results → Continuous vllm-exl3 and ExLlamaV3 kernel development → SixCat quality, speed, concurrency, and long-context benchmarks → Reproducible one-Spark recipes the community can run themselves → Faster debugging when new models or inference engines break It would turn a lot of: “I built it, but I need someone with a Spark to test it” into: “I built it, tested it, documented it, and released the recipe.” 𝗠𝗬 𝗣𝗨𝗕𝗟𝗜𝗖 𝗪𝗢𝗥𝗞 𝗦𝗧𝗔𝗬𝗦 𝗣𝗨𝗕𝗟𝗜𝗖 This is not an access fee, preorder, or paywall. The models, recipes, benchmarks, fixes, and tools will remain public. This is completely optional support for the hardware behind that work. I’m planning to place the order within 48 hours and will personally cover whatever the fundraiser does not. For complete transparency, I’ll publish: → The purchase receipt with private details removed → Photos of the actual DGX Spark → Total community contributions → Platform and payment fees → Final purchase cost → The amount I contributed myself To everyone who has loaned me hardware, funded compute, tested a build, submitted a PR, shared my work, or simply encouraged me: thank you. I genuinely could not have pushed this work as far as I have without this community. A contribution would mean a lot. A repost helps just as much. ❤️ ko-fi.com/victorangelcruz/go…
31
33
154
55,615
It’s been about 15 hours since I posted my DGX Spark fund and I’m absolutely blown away by how quickly the community has come together for me! I didn’t have courage initially to post this fund as I really didn’t want to ask or be a beggar, I thought coming into this that I would be worthy of such love and donation from this community! 😩 Thank you all and God bless you for any reposts and every single penny invested into my work! I will continue to push forward 🚀 So far I’m at $3,410.00 donated - 68.2% of the $5,000 goal. $1,590 left!! 🤯🤯 - Ko-fi: $1,910 (38.2%) - BTC: $1,500 (30.0%) 🟩🟩🟩⬜⬜⬜⬜⬜⬜⬜ 38% of goal on Ko-fi! Donation link: ko-fi.com/victorangelcruz/go… My BTC address: 3GT8KQuh9xagVbFsGjHSvaNAUW5bXsty3i
A lot of people in this community were genuinely shocked when they learned this: I do not own a DGX Spark. Much of the Spark-specific work I’ve released, including day-zero EXL3 and GGUF quants, vllm-exl3 development, custom kernels, SixCat evaluations, and reproducible serving recipes, has only been possible because incredible people in this community have loaned me access to their own hardware and clusters. I’m deeply grateful for that support. But borrowed access is temporary and unpredictable. I often have a new model, fix, or benchmark ready, then have to wait until someone else’s machine is available before I can validate and publish it properly. So I’m raising **$5,000 toward a DGX Spark of my own.** 🙏 𝗪𝗛𝗔𝗧 𝗔 𝗗𝗘𝗗𝗜𝗖𝗔𝗧𝗘𝗗 𝗦𝗣𝗔𝗥𝗞 𝗪𝗢𝗨𝗟𝗗 𝗘𝗡𝗔𝗕𝗟𝗘 → More true day-zero quantized model releases → Immediate hardware validation instead of projected results → Continuous vllm-exl3 and ExLlamaV3 kernel development → SixCat quality, speed, concurrency, and long-context benchmarks → Reproducible one-Spark recipes the community can run themselves → Faster debugging when new models or inference engines break It would turn a lot of: “I built it, but I need someone with a Spark to test it” into: “I built it, tested it, documented it, and released the recipe.” 𝗠𝗬 𝗣𝗨𝗕𝗟𝗜𝗖 𝗪𝗢𝗥𝗞 𝗦𝗧𝗔𝗬𝗦 𝗣𝗨𝗕𝗟𝗜𝗖 This is not an access fee, preorder, or paywall. The models, recipes, benchmarks, fixes, and tools will remain public. This is completely optional support for the hardware behind that work. I’m planning to place the order within 48 hours and will personally cover whatever the fundraiser does not. For complete transparency, I’ll publish: → The purchase receipt with private details removed → Photos of the actual DGX Spark → Total community contributions → Platform and payment fees → Final purchase cost → The amount I contributed myself To everyone who has loaned me hardware, funded compute, tested a build, submitted a PR, shared my work, or simply encouraged me: thank you. I genuinely could not have pushed this work as far as I have without this community. A contribution would mean a lot. A repost helps just as much. ❤️ ko-fi.com/victorangelcruz/go…
4
8
61
2,838
DeepSeek-V4.1-Flash EXL3 is finally alive across 4× DGX Sparks. 🔥 I built the quality-first 4.75 bpw SAGE quant. @cfontes and @Blackwellboy spent more than a week, 14 PRs, dozens of commits and 40+ bring-up iterations making it actually run. I don’t own one Spark, let alone four. This is what open-source collaboration looks like. 𝗧𝗛𝗘 𝗥𝗘𝗖𝗘𝗜𝗣𝗧𝗦 Single-stream decode: 30.2 tok/s prose 33.4 tok/s code 33.37 tok/s best warm sustained cell Aggregate at c=1 / 2 / 4 / 8: 27.8 / 43.9 / 57.7 / 61.5 tok/s KV budget: 3.97M tokens Context: full 1,048,576 tokens Every figure is a sustained 30-second llm-inference-bench cell, not a burst. 𝗪𝗛𝗔𝗧 𝗖𝗛𝗥𝗜𝗦 𝗛𝗔𝗗 𝗧𝗢 𝗦𝗢𝗟𝗩𝗘 Across the recipe and vllm-exl3: 14 PRs 12 merged Dozens of commits He: → got the GB10/aarch64 build and four-node RoCE cluster working → recovered 1,243 missing non-expert backbone tensors that had been producing deterministic nonsense → fixed EP-aware loading for the full 233 GiB expert bank → repaired multiple SM121 sparse-attention and page-geometry blockers → cut the TP4 cold load from ~8.7 minutes to ~5 minutes → built mixed-K fused MoE and CUDA-graph-safe padded kernels → caught a wrong EXL3 codebook path that produced plausible-looking but incorrect output → discovered the DSpark loader left three ranks with ZERO draft experts, meaning only 25% of the draft expert weights were loaded → fixed disk-backed Engram replay and cold NVMe row loading → unblocked full 1M context One loader fix alone recovered 9–20% decode. His Engram I/O work improved novel-text steps by 27.4%. This was not “run a script and benchmark it.” It was a week of tracing silent correctness bugs, writing kernels, rebuilding four machines, profiling stalls, validating output and refusing to call it done until the receipts were real. 𝗪𝗛𝗬 𝗘𝗫𝗟𝟯 𝗢𝗩𝗘𝗥 𝗡𝗩𝗙𝗣𝟰? NVFP4 is an excellent throughput-first Blackwell format, but it optimizes a different target. NVIDIA’s V4.1 release serves the routed experts as W4A4: 4-bit weights and 4-bit activations with fine-grained block scaling. Its expert-weight conversion is effectively lossless relative to DeepSeek’s shipped MXFP4 experts, so the quality argument is not simply “4.75 bits beats 4 bits.” My 4.75 bpw SAGE-EXL3 pack is built for the highest fidelity I could preserve within a four-Spark budget: → QTIP-derived, Hessian-aware trellis quantization → Mixed K3–K8 instead of one flat precision everywhere → More bits allocated to the tensors that are most sensitive → No requirement to quantize the expert activations to FP4 → Engram retained in native FP8 → Attention, shared experts, DSpark/MTP, vision and the output head preserved rather than wholesale-requantized EXL3 therefore has a higher fidelity ceiling: it spends a larger and more flexible precision budget on the expert bank while avoiding the additional FP4 activation rounding present in an NVFP4 W4A4 execution path. The protected parts of the architecture also remain at their source precision. That is why I expect this 4.75 bpw pack to better preserve difficult reasoning, coding, vision and long-horizon agent behavior. To be completely precise, the controlled EXL3-vs-NVFP4 KLD and benchmark suite is still pending, so this is the technical quality expectation and design goal, not yet a claimed measured victory. I created the highest-quality four-Spark quant I could make, but without Chris’s hardware access, patience and engineering, it would still be a set of weights I had no physical way to validate. People like him are why this community moves forward. One person creates the quant. Another brings the hardware and runtime expertise. Everyone gets the result. Thank you, Chris and Aaron, Seriously. 🙏 There is still room to improve the drafter, but this milestone is real. Quant: huggingface.co/vcruz305/DSV4… Recipe + receipts: github.com/vcruz305/DeepSeek… vllm-exl3: github.com/vcruz305/vllm-exl…
We got @DeepSeek-V4.1-Flash running in EXL3 on 4× Sparks! I was happy to colab with Cruz (@ViC305) and get his vllm-exl3 and his DeepSeek-V4.1 recipe working on TP=4! Single-stream decode: 30.2 tok/s prose, 33.4 tok/s on code, with a best warm 33.37 tok/s Aggregate: 27.8 / 43.9 / 57.7 / 61.5 tok/s at c=1/2/4/8. As you can see, there is opportunity for growth here, but it is going to take a week or so of addressing some drafter issues that might not get me much, so I am switching gears to another project for now, but might come back to this later. KV budget: 3.97m tokens Context: full 1m Every figure is a 30-second sustained cell from llm-inference-bench, not a burst. Get it here! github.com/vcruz305/DeepSeek…
4
6
31
3,485
One 36B MoE. Six EXL3 releases. GPUs from 8 GB to 96 GB. 🔥 K2-Horizon-MoVA-36B-A4B is now local for almost everyone!! 👇🏼 2.5 bpw: 13.27 GB with 83.7% top-1 agreement 8.0 bpw: 38.07 GB with 96.43% agreement, at roughly half the BF16 weight footprint This is exactly why I use EXL3 when the goal is maximum model quality per GB. 𝗞𝟮 𝗛𝗢𝗥𝗜𝗭𝗢𝗡 𝗜𝗦 𝗔 𝗩𝗘𝗥𝗬 𝗜𝗡𝗧𝗘𝗥𝗘𝗦𝗧𝗜𝗡𝗚 𝗠𝗢𝗗𝗘𝗟 It stores 36B parameters but activates only about 4B per token through a combination of: → Mixture-of-Experts feed-forward routing → Mixture-of-Values attention → 512K native context → fully open weights, training data, logs and intermediate checkpoints Against Qwen3.6-35B-A3B, IFM’s published BF16 results win 6 of the 9 listed comparison rows, including agentic tool use, Terminal-Bench, SciCode, HLE, CritPt and non-hallucination. Now I’ve released the complete EXL3 ladder for local GPUs. 𝗣𝗜𝗖𝗞 𝗬𝗢𝗨𝗥 𝗚𝗣𝗨 48 GB+ maximum quality: 8.00 bpw · 38.07 GB · 96.43% top-1 48 GB with more KV/context headroom: 6.50 bpw · 31.34 GB · 90.68% 32 GB: 5.00 bpw · 24.56 GB · 85.83% 24 GB: 4.00 bpw · 20.05 GB · 84.81% 16 GB: 2.50 bpw · 13.27 GB · 83.71% 8 GB + expert offload: 2.00 bpw · 11.01 GB · 81.66% Every pack was scored against the original BF16 model over 10,240 held-out token positions before upload. Top-1 agreement means both models selected the SAME most-likely next token after receiving the same preceding text. KLD measures the full probability distribution. Lower means the quant is tracking BF16 more closely. For perspective, unquantized BF16 running through the same ExLlamaV3 implementation reaches 97.71% agreement with IFM’s reference. The 8-bit EXL3 pack reaches 96.43%. That is extremely close to the practical ceiling while cutting the model to 38 GB. 𝗡𝗢𝗧 𝗔 𝗙𝗟𝗔𝗧 𝗤𝗨𝗔𝗡𝗧 These were built with my internal SAGE mixed-precision workflow for EXL3. Instead of applying the same precision blindly across the entire model, the allocation is adapted across the model body while keeping the output head at 6-bit. The result is a ladder designed around actual GPU capacities instead of arbitrary bitrate labels. 𝗡𝗘𝗫𝗧 𝗨𝗣 I’m actively working on: → RTX PRO 6000 benchmarks → decode and prefill measurements → MTP/speculative decoding → MoVA + MoE kernel improvements → context and memory testing → a full reproducible serving recipe Speed receipts and improvements are coming next. Weights + quality results: huggingface.co/vcruz305/K2-H… Recipe: github.com/vcruz305/K2-Horiz…
K2-Horizon-36B-A4B scores 25 on the Artificial Analysis Intelligence Index, matching models with over 20× the total parameters while using 4B active parameters per token. These capabilities come from our new architecture MoVA (Mixture-of-Value Attention), which incorporates MoE-based sparsity into the compute of value vectors in multi-head attention. It opens a second axis for scaling sparsity in an LLM, beyond MoE in the FFN module. Importantly, MoVA enjoys the following advantages: • Simple and compatible with efficient attention algorithms, such as flash attention, GQA, and sparse attention • No additional KV cache cost comparing to standard GQA K2-Horizon-36B-A4B available at: huggingface.co/IFM/K2-Horizo…
7
8
71
14,485
Ooooo Claude Opus 5.5 loading up for a release!! Let’s go Anthropic make us proud !!!
New model: claude-opus-5-5 Maker: Anthropic Detected in Anthropic API models.
1
1
6
926
Great post on what models to use with what hardware! Extremely needed when you have people like me posting about Spark models all day 😂 I’ll be following up this post with a new model ladder that will have quantizations for every GPU size starting at the 96GB rtx6000 pro down to the 8GB little homies 🫡 stay tuned!!
So You Just Bought the Hardware. Now What? I keep getting asked some version of: “I have X hardware but what model should I actually run on it?” There are honestly just the setups id actually start with right now. Weighing the reliability pretty heavily not just benchmark scores. If ive actually abused it on my own hardware for hours without it failing, that counts for a lot.. 1× DGX Spark → Qwen3.8 Flash Next Probably my favourite all round starting point on one Spark right now. Huge model for one box pretty gooood agent/tool potential..also there are multiple solid recipes for it. I put the Unsloth UD-Q2_K_XL build through a full 6 hour soak too: 9,433 requests, 0 failures no crash. So this one isnt just me liking the benchmark numbers. 2× DGX Sparks → MiMo 2.6 / GLM-5.3 Flash This is where things get really interesting...so after testing MiMo properly Id actually put it right at the front now. Tonys TP2 recipe reproduced almost bang on for me it hit 159.98 tok/s at C6, then survived a full 6 hour mixed workload with 65,226 requests and 0 restarts... 0 OOM and 0 NCCL...yeah. let that soak in, haha, sorry. GLM-5.3 is still right beside it. The EXL3 work on two Sparks has gotten really good and the recipes are much more mature than they were i guess not too long ago. so If I wanted the more established path Id probably lean GLM. If I wanted to play with what impressed me most this week hands down MiMo. 3× DGX Sparks → GLM-5.3 Flash Same answer. Dont change model jst because you bought another Spark 😂 TP3 gives it more room and this is one of those setups where the third box actually makes sense...even though i find a hard time commiting to anything past TP2, its honeslty a sweet spot. If you can get it to run well and you save a spark...why wouldnt you. 4× DGX Sparks → DeepSeek V4.1 Flash This is where I’d move to the bigger DS setup. TP4 much more room to breathe and current recipes have actually demonstrated the full 1M context instead of just putting 1000000 in a config file and pretending shes as stable as my partner. Personally dedicating 4 Sparks to one model instead of running two really good TP2 lanes is really not my thing at the moment. RTX 5090 32GB → Qwen3.8-27B This is king. Fight me. The amount of performance we've maaged to drag out of this model on a 5090 with autoround/exl3/dflash stuff is insane. RTX 3090 24GB → Qwen3.8-27B This card and its ability to run THIS model is just unbelievable. Limits were pushed, boundaries were broken. And before someone starts a war in the replies: these are not the objectively best models in existence theyre the ones I think make the most sense for the hardware. thats a completely different question. These guys have done a ridiculous amount of work on the Spark side, and @ViC305 @turboderp_ and @Tech2Wild have some really interesting EXL3/other builds Id use depending on whether I cared more about quality..context, speed or just seeing how far I could abuse the hardware. Im just linking the builders below instead of dumping 19 individual recipes on you. Go shopping based on whether you care more about quality, context, speed, or seeing how badly you can abuse the hardware. Vic : github.com/vcruz305 Turbo : huggingface.co/turboderp Tony : github.com/tonyd2wild Unsloth : github.com/unslothai
2
2
11
1,285
Qwen3.8-Flash-Next: EXL3 ~600 vs NVFP4 ~150 tok/s reported concurrent decode. 🤯 Both on DGX Spark GB10. The interesting part is what the engines actually do with those compressed weights. This will be my last informational post on Qwen3.8-Flash-Next as im cooking up something new for you all! 👨🏻‍🍳 3.05 bpw EXL3 through my ExLlamaV3 fork. NVFP4 through vLLM. 𝗧𝗛𝗘 𝗗𝗘𝗖𝗢𝗗𝗘 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗖𝗘 Balanced workload, EXL3 / NVFP4: Single-request decode: 78.3 / 31.0 tok/s Per-request decode at 8 requests in flight: 75.9 / 18.9 tok/s Time per output token at that load: 13.2 / 53.4 ms Reported concurrent decode: ~600 / ~150 tok/s The longer 512-token output workload reports: 569 / 176 tok/s concurrent decode. Those concurrent figures are the harness’s reported metric, not wall-clock server throughput. Overlapping decode windows still need verification before treating them as eight simultaneously generating streams. The individual-request speed difference is clear. 𝗪𝗛𝗬 𝗧𝗛𝗘 𝗘𝗫𝗟𝟯 𝗣𝗔𝗧𝗛 𝗜𝗦 𝗦𝗢 𝗙𝗔𝗦𝗧 EXL3 uses trellis-coded weights. ExLlamaV3’s specialized decode kernels reconstruct those weights as they multiply, with paths designed for small batches. NVFP4 uses scaled 4-bit floating-point values. Blackwell can accelerate that math directly, but the actual vLLM backend and workload shape matter. A format’s peak Tensor Core throughput is not the same thing as low-latency token generation. My native recipe also has several measured optimizations working together: • Wider fused-MoE decode tiles selected for GB10. • INT8 mixer-weight storage, worth 7–13% decode in separate A/B tests. • Five-token MTP with dynamic stopping, amortizing target-model work across accepted tokens. • A smaller draft-head projection, with the full head retained for verification. • 8-bit KV to reduce attention-cache traffic. • Fewer host synchronizations and launch threads pinned to the big CPU cores. One surprising detail: I turned OFF the INT8-activation GEMV path because FP16 GEMV was faster on this Spark. That is different from the INT8 mixer weights, which DID help. This is why tuning the actual execution path matters more than picking the lowest-sounding datatype. 𝗙𝗔𝗦𝗧 𝗧𝗢 𝗟𝗢𝗔𝗗, 𝗧𝗢𝗢 The native EXL3 recipe records a ~47-second cold model load vs 200+ for NVFP4. Packed weights load directly, while the large n-gram table stays file-backed rather than being copied wholesale into GPU allocations. That avoids making the entire table resident before serving. Cold lookups can still pay for page-ins later. 𝗧𝗛𝗘 𝗪𝗛𝗢𝗟𝗘 𝗥𝗘𝗤𝗨𝗘𝗦𝗧 𝗦𝗧𝗜𝗟𝗟 𝗖𝗢𝗨𝗡𝗧𝗦 One 96-token answer: NVFP4 starts sooner: 416 ms vs 583 ms TTFT. EXL3 finishes sooner: 1.80 s vs 3.55 s. At 8 requests in flight, however, whole-workload aggregate favors NVFP4/vLLM: Balanced: 64.8 vs 52.0 tok/s 512-token outputs: 82.1 vs 66.4 tok/s Faster individual streaming and higher server throughput are different wins. Same GB10 hardware class and sampling settings, but different formats, engines and execution configurations. This is not an isolated kernel A/B or a model-quality comparison. For the measured individual-request workloads, my EXL3 + ExLlamaV3 deployment is substantially faster. The recipe shows the engineering behind it, not just the biggest number. Credit to turboderp / ExLlamaV3 for the format, engine and kernels. Recipe + measurements: github.com/vcruz305/Qwen3.8-… Native fork + optimizations: github.com/vcruz305/exllamav…
9
12
101
7,660
Cruz retweeted
Replying to @ViC305
@ViC305 Very Fast Qwen 3.8 Flash Next EXL3 on 1 spark. Great work
2
4
10
510
Qwen3.8-Flash-Next EXL3 crossed 100 tok/s on ONE DGX Spark. 🔥 102.6 tok/s on repetitive code in a community retest of my native ExLlamaV3 recipe. And the part I’m happiest about? My ~79 tok/s code result reproduced at 79.5 on their Spark. 𝗧𝗛𝗘𝗜𝗥 𝗕𝗘𝗡𝗖𝗛𝗠𝗔𝗥𝗞 𝗥𝗘𝗦𝗨𝗟𝗧𝗦 → 102.6 tok/s: repetitive Python code → 79.5 tok/s: my 400-token code workload → 71.7 tok/s: a 600-token prose explanation → 66.2 tok/s: JSON catalog Thinking off, greedy sampling, better of two runs. These are decode rates, not total request throughput. The 100+ result is a highly predictable workload where speculation shines. It is not a blanket “every coding task runs at 100 tok/s” claim. 𝗟𝗢𝗡𝗚 𝗖𝗢𝗡𝗧𝗘𝗫𝗧 𝗛𝗘𝗟𝗗 𝗨𝗣 𝗧𝗢𝗢 15/15 needle-retrieval checks passed, including one buried 243,635 tokens deep. 1,000+ tok/s prefill at that depth. Clean tool calls. An 80K tool loop also ran on the setup. MTP depth 5, dynamic stopping, 8-bit KV and int8 mixers are doing real work here. This is why I publish the recipes. Someone else reproduces the results, tests workloads I haven’t, and sends improvements back. Their fixes were merged the same day. Huge thanks for putting it through the full grid, and credit to turboderp for ExLlamaV3. One small box. A lot more room to experiment. Thank you so much to @yume_arasaki for the time invested and hard work on this benchmark!! Recipe + native runtime instructions: github.com/vcruz305/Qwen3.8-…
I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX Spark. Cruz shipped a custom ExLlamaV3 fork and a native engine path. I ran the same model on the same Spark: 102.6 max, 80 tok/s on code, 70 tok/s on prose. My previous benchmark and community recipes was 48.6 tok/s. I ran my full grid on it so you don't have to. This might now be the best way to run this model, without breaking the bank on one DGX Spark. THE QUANT EXL3 at 3.05 bits per weight. Trellis coding packs information more efficiently than plain rounding, so 3 EXL3 bits should land near 4.5 conventional bits. On paper that sits above FP4. I wasn't able to verify this, subjectively, running this on my harness feels at least on par FP4, and definetely not worse. A single source says NVFP4 is better on longer context. THE RECIPE One line for the DD crowd: vcruz305/exllamav3 at 523ecd3, draft depth five, dynamic stop 0.6, 8-bit KV, int8 mixers with exact fp16 reduction order, big cores pinned. Every flag in that ladder earned its tok/s. The mixers shipped fp16 inside a 3-bit model, and storing them int8 is worth 7 to 13 percent decode on its own. The 8-bit KV adds 3 tok/s at 4k and 7 at 240k, where twelve full-attention layers stream the whole cache every step. THE NUMBERS All measured on my Spark this week. Thinking off, T=0, empty KV, decode after first token, better of two reps. A speed record from a bench that cannot multiply means nothing: 17 × 19 → 323 runs first, every grid. Code (50 identical Python clamps): 102.6 tok/s. Counting 1 to 200: 107.1, with 99 percent of drafts accepted. Nothing is more predictable than numbers in order. Prose (explain a hash map, 600 tokens): 71.7, two reps at 70.5 and 71.7. JSON catalog: 66.2. Cruz's published number on the same fork: 79 tok/s code. Mine is 102.6 because repetitive code clamps accept nearly every draft. His 53 prose is a 350-word story. My 71.7 is a 600-token explanation. Same SHA, different jobs. Both stand. Sanity check: a Cruz-style greedy 400-token code job runs 79.5 on my card, 74 percent draft acceptance. His 79 reproduced. CONTEXT Needle at 5/50/95 depth, three codes: 15 for 15, one buried 243,635 tokens deep. Prefill above 1,000 tok/s at that depth. Decode at depth (256-token jobs, KV resident): 8k 51.5, 32k 55.6, 64k 67.4, 110k 66.0, 240k 55.8. Decode rises into the mid 60s at 64k before settling. No cliff. With the cache already full: count 112.9 at 32k, 82.0 at 128k. Prose 52.2 at 32k, 49.0 at 128k. Code 111.1 at 32k, 69.0 at 128k. Tools called at every depth, no XML leaking into content. AGENT LANE Tools on at 35k resident: proper tool calls returned, 58 and 57 tokens, prefill about 980 tok/s. Trust check: the first pass printed agent decode up to 386,000 tok/s. Pretty kickass. Tools off, forced 2,048-token HTML output: 86.8 and 88.3 tok/s engine rate, 86-87 percent draft acceptance. Wall clock lands near 35 tok/s because a 35k prefill puts 36 seconds behind every first token. Decode was never the agent bottleneck on this box. Prefill is. The research client ran an 80k tool loop on this serve and lived. ENERGY The prose cell: 0.85 joules per token on one GPU rail. At that rate: 6.2 million tokens a day for about 27 cents of electricity. The same tokens on Grok 4.6 output pricing are $37 a day. The box still costs a car. THE HONEST LIMITS The native engine runs one job at a time. If you need concurrent streams, Cruz ships a vLLM lane alongside it. Different engine, lower single-stream speed, real concurrency. Pick per workload. MY READ This recipe is the killer app for a single Spark. One quiet box, frontier weights, 100 tok/s code, 71 prose, a 262k window that holds, tools clean at depth. If you own one Spark, this is the loadout. Sent Cruz a PR with light fixes from this grid. Merged same day. 100 tok/s on code from one small box was not supposed to happen this year. Benchmark, sources and recipe in reply 👇
16
17
120
11,719
About 2 weeks or so later since hitting 1,000 and I’ve already surpassed 2,000 followers!! 🤯 Extremely grateful to God for the will to keep persisting for the community! Im not in this for the glory or fame, I’m just a one man team that does his own quantizations and recipes trying to push the bar forward for local inference! 🫡 I’m also extremely blessed to have awesome people like @Blackwellboy, @cfontes, @aijoey, @WescheNex1q and @justinn_builds in my court who are willing to put their projects on hold to work on validation, benchmarks and recipes! So maybe I’m not totally a one man team with these guys with me! With added on blessings from @WescheNex1q, @Blackwellboy, @MarkusEicher70, @CountStaculaAI and @GGWoodsman, @Blackfrost_AI and @RevAndrei who have been gracious enough to give me spark and RTX6000 access to continue my work! (as I don’t own one…yet!) Here’s to 2,000 more! I promise to never stop pumping out that 🔥 for all of you! God bless you all who have followed me and helped me and God bless all the nay-sayers and haters as well for I love you all equally!! All the glory goes to Jesus! ✝️
12
4
45
2,496