I run local AI at home on an RTX 5090, two RTX 5060 Ti and four DGX Sparks, and rank what runs best on the machine you own. Every number with receipts.

Germany
What I'm building: one plain ranking of which AI runs best on the machine you already own. 16 GB card, 32 GB, one DGX Spark, up to four Sparks linked in a ring. Same prompts on every class: a PDF, images, a voice memo, a bug fix, a game. Every number with receipts. Today: 30+ runs across six machines. Results start landing here. I run it at home so you don't have to guess.
2
1
8
1,131
I saw this post by @jayleaton yesterday and it got me curious: Qwen-Image on TensorFold's NVFP4 kernels, 33.0 s down to 8.2 s per image on a Windows 11 + RTX 5070 Ti. So I ported his repo to Linux and ran it on my RTX 5090. The repo is jayleaton/qwen-image21-tensorfold-rtx, HEAD at commit f918455 (no tagged release). It runs Qwen-Image on TensorFold by @ashxhart inside ComfyUI. The test: RTX 5090 at 450 W, 1024x1024, 25 steps euler, same prompts, same seeds. Seconds per image, median of 5, first run dropped as warmup: GGUF Q6_K (ComfyUI's usual loader): 13.87 s TensorFold FP8: 6.94 s (2.0x faster) TensorFold NVFP4: 4.53 s (3.1x faster) The price is memory. VRAM peak: GGUF 15,750 MiB, TensorFold 22,940 MiB (NVFP4) and 23,036 MiB (FP8). Power draw mean 359 W on GGUF, 302 W on NVFP4, 450 W limit untouched, max 64 °C. The images are not identical. With the same seed, every path renders a different image: all 36 files are unique. Composition and text stay close, background details drift. He sees the same: his own quality gate (DINO 0.95) fails on both quantized paths, 0.908 NVFP4 and 0.920 FP8, LPIPS passed. I could not rerun DINO/LPIPS, no local weights. What differs from his setup: the Q6_K file was quantized locally from the bf16 checkpoint (patched llama.cpp image quantizer), the engine converted from that same file, and NVFP4 ran without his calibration pass. His numbers come from Windows 11 + RTX 5070 Ti, not directly comparable to my Linux + RTX 5090: he reports 33.0 s GGUF, 16.7-17.8 s FP8, 8.0-8.4 s NVFP4. All images shown are AI-generated. Shoutout to @jayleaton for building this and sharing it, and to @ashxhart for the TensorFold kernels. The engine is real, nice work. Repo: jayleaton/qwen-image21-tensorfold-rtx @f918455
Qwen-Image 2.1 on an RTX 5070 Ti 16 GB: 33s per 1024² image, down to 8. @ashxhart TensorFold's NVFP4 kernels, dropped into my normal ComfyUI as a loader node. Same workflow, same text encoder, sampler and VAE. 1024²: 33.0s to 8.2s (4.0x) 992×1216: 35.7s to 9.4s (3.8x) 480×608: 16.1s to 1.9s (8.4x) The catch: it's not pixel-identical. Same subject and clean images, sign text and hands are still right, but fine details drift slightly. To be fair, I have used this to generate hundreds of images and had a similar result even before the optimisation pass. FP8 stays closer to the original for strict editing and still gets a 2x speed boost. Side by side below. Judge for yourself. The recipe will be in the comments soon. DW, I am still cooking something for the DGX Sparks at the same time.
1
1
171
Part 2 of my local AI test is live. Part 1 was a few days ago: 5 setups on 4 memory classes, 10 app prompts, 45 minutes each. It's still on the page. Part 2: 9 setups, same 4 classes, same 10 tasks. New: kits and recipes from @MiaAI_lab and @ashxhart, and the 3 biggest tasks get up to 150 minutes. 16 GB (RTX 5060 Ti) · Qwen3.8-27B Mia's exllamav3 kit and my llama.cpp 32 GB (RTX 5090) · Qwen3.8-27B Mia's vLLM recipe and Ash's TensorFold 128 GB (GB10) · Qwen3.8-Flash-Next Mia's TensorFold recipe and Mia's vLLM kit 256 GB (2× GB10) · GLM-5.3-Flash Mia's vLLM, Mia's TensorFold, Ash's TensorFold Every app they built is clickable, task by task. Details in the replies. nipale-ai.github.io/local-ai…
5
5
315
256 GB (two GB10 boxes per cell), model GLM-5.3-Flash: Mia's vLLM kit: vLLM, EXL3 4bpw 54.43 tok/s, TTFT 2.62 s warm Mia's TensorFold kit: TensorFold 0.6.0, EXL3 4bpw 89.45 tok/s, TTFT 1.13 s warm @ashxhart's recipe: TensorFold 0.6.0, MLX 4bit 55.66 tok/s, TTFT 1.19 s warm tok/s = output tokens per second per request, median. TTFT = p50, warm = prefix-cache hit. Apps: nipale-ai.github.io/local-ai…
39
Replying to @MiaAI_lab @ashxhart
128 GB (one GB10 box each), model Qwen3.8-Flash-Next, both are Mia's recipes: TensorFold 0.3.6.3, MLX 4bit 105.5 tok/s, TTFT 4.99 s (first request of a task), 0.82 s later vLLM kit, NVFP4 72.8 tok/s, TTFT 2.12 s (first request of a task), 1.21 s later tok/s = output tokens per second per request, median. TTFT = p50. Apps: nipale-ai.github.io/local-ai…
20
Replying to @MiaAI_lab
32 GB (RTX 5090), model Qwen3.8-27B on both sides: Mia's recipe: vLLM 0.27.1, NVFP4 75.43 tok/s, TTFT 20.39 s (prefix cache off, every request cold) @ashxhart's TensorFold 0.6.0, MLX 4bit 352.28 tok/s, TTFT 30.01 s cold, 0.84 s warm tok/s = output tokens per second per request, median. TTFT = p50. Apps: nipale-ai.github.io/local-ai…
1
32
Replying to @MiaAI_lab @ashxhart
16 GB (RTX 5060 Ti), model Qwen3.8-27B on both sides: Mia's kit: exllamav3 1.4.4, EXL3 2.5 bpw 56.34 tok/s, TTFT 3.98 s My setup: llama.cpp b11151, UD-Q3_K_XL (about 3.8 bpw) 39.71 tok/s, TTFT 2.51 s tok/s = output tokens per second per request, median. TTFT = p50 over all requests. Apps: nipale-ai.github.io/local-ai…
26
Excited that Mia is getting into local image generation! I already have data from my September test: Qwen-Image 2.1 vs FLUX.2 [dev], 8 prompts, 2 seeds, 64 images in total, on an RTX 5090 (32 GB) and an RTX 5060 Ti (16 GB). Happy to share the images and prompts for your series. The image shows prompt A1 (sakura-911), one of the 8, seed 4210, same seed in all four cells. My rating for the images shown: Qwen 9 vs FLUX 5 on the 5090, 9 vs 4 on the 5060 Ti. On the other seed of this prompt (4211) FLUX got 7 and 6, Qwen 9 and 9. Across all 64 images my rating averages Qwen 7.1 vs FLUX 4.5 on the 5090 and 7.1 vs 4.4 on the 5060 Ti. Per image: 86 s vs 172 s on the 5090, 490 s vs 722 s on the 5060 Ti. Each card ran its own build, so builds differ: Qwen bf16 (5090) and int8 (5060 Ti), FLUX GGUF Q6_K and Q3_K_S, with other text encoders too. FLUX does not fit into 16 GB even as Q3, so on the 5060 Ti it streams its weights from NVMe. Steps: Qwen 40, FLUX 20. Going from the 5090 to the 5060 Ti setup moved my rating by about 0.1 points, going from my Qwen to my FLUX setup by 2.6 (on the 5090) and 2.7 (on the 5060 Ti). My own rating, one person. All images AI-generated. All 64 images side by side, full prompts, timings and my rating table are here: nipale-ai.github.io/qwen-ima…
Getting into local AI image generation! Starting with two recent open models: Ming-Image-0.1-Design and Qwen Image 2.1. This is the first post in a series of comparisons I'm going to do between them. I ran both on my RTX 5090. Same prompt, same seed. Both did a great job. Ming's image looks more cinematic. Qwen's is clean and simple. Two different styles, both usable imo. More tests coming in the next few days. Model links and full prompt below.
3
311
Window for prompt speed: what turboderp's EXL3 packs change on a 32 GB card. Same RTX 5090 (power limit 450 W), same TensorFold 0.6.2 (@ashxhart; 0.6.3 is out), same Qwen3.8-27B + z-lab DFlash2 drafter, greedy. nvidia's NVFP4 checkpoint next to turboderp's EXL3 packs at 2.50bpw and 2.00bpw. Window = max prompt+reply per request, set at boot (NVFP4 · EXL3-2.5 · EXL3-2.0, tokens): parallel 1: 84,542 · 202,542 · 224,170 parallel 4: 29,180 · 143,085 · 163,959 parallel 8: 3,453 · 112,691 · 132,723 parallel 16 (patched build only): refused at boot (window 0) · 25,599 · 44,122 Drafted decode, one request, server at parallel 8, Ash's bench_concurrent with 256 tokens (median of 3 boots, 3 runs each): code prompt NVFP4 293.8 tok/s, EXL3-2.5 273.3 tok/s; chat prompt 204.0 tok/s, 138.7 tok/s. Cold prefill on a 2k prompt, same servers (median of 18 runs): 8,031 tok/s, 2,543 tok/s. EXL3-2.0, same servers, 2 boots, one patched: decode 362.1 tok/s code, 209.3 tok/s chat (median of 2 boots, 3 runs each), prefill 2,540 tok/s (median of 12 runs). So both EXL3 packs pay for their window in prompt speed. Decode is mixed: EXL3-2.0 was faster than NVFP4 on both prompts, EXL3-2.5 slower. Token-exact: drafts matched serial in all 16 suite arms. Needle found in all 16 runs (31,533-token prompts on the EXL3 arms, 1,008-token prompts on NVFP4, because at parallel 8 its window is 3,453). Honest limits: I measured window and speed, not quality. The packs sit far apart in bits per weight (NVFP4 ~4.5 effective on its FP4 layers, FP8 and bf16 around them; EXL3 2.50 and 2.00), so more window does not mean the same answer quality. TensorFold boots EXL3 packs as experimental. One run day, one card. My 16 GB tight-staging patch (earlier post) ran in the patched arms: in the A B B A A B pairs it changed neither window nor speed (medians within 0.4 %). NVFP4 boot peak, highest of 6 servers each on the 1 s sampler: release 23.4 GiB, patched 20.9 GiB. The patch is not in TensorFold.
2
5
663
Niklas retweeted
The funny part is all the big tech people already know who Mia is. We just dont really care. What are we doing being parasocial 😭 if you like someone's content follow it if you dont then dont. Shi just seems weird to me ngl. What if we all just focus on what we got going on haha
1
3
46
4,128
If you spend a significant amount of time on this platform and want to grow your account, this is definitely an investment worth making. The book answers exactly the questions you need answered. No fluff, straight to the point. Of course I bought it myself.
To celebrate this unusual event, I'll be offering 25% off my book for a limited time. Use stefanmaier-25 in the coupon code. The book is literally how I grew my account! Get it here: mia-ai.net/book
1
5
1,235
Niklas retweeted
I have been hearing the exact same takes every single day for over a year now. Let me break it down for you again: 1. ​Local AI running on setups like DGX Sparks or Mac Studio is already hitting faster tokens per second than most Cloud AI. 2. ​Back when $20 subs were standard, people claimed local rigs cost years of subscription fees. Now people burn through $500 monthly plans and still hit usage caps. It will only get more expensive. ​3. People who bought their hardware a year ago, or even just a few weeks ago, are already seeing up to a 200% return simply from hardware price jumps. 4. ​Local AI never gets nerfed overnight, and it never goes down because of server outages. 5. ​Absolute raw intelligence is slightly behind top cloud models, but it easily matches last gen flagships. That is more than enough for most real work. On top of that, you get zero censorship and full privacy that Cloud AI cannot offer.
Replying to @jun_song
There is a hard limit on what a DGX Spark can do. The machines are slow as hell. You could rent years of cloud GPUs for a fraction of the price and at multiples the speed. Today, a spark is useless. In 3 years? A DGX spark is worth less than an old laptop.
28
35
317
18,599
I haven't seen a 16GB card yet, so I patched @ashxhart's TensorFold 0.6.2 so Qwen3.8-27B (EXL3) boots and serves on a 16 GB card, an RTX 5060 Ti. Opt-in env flags, defaults untouched, not published. What blocked it: admission billed staging at three times the largest tensor group, and on EXL3 packs that group is the bf16 embedding table. My patch bills a tighter, EXL3-only estimate per group instead, and packs the drafter weights in chunks. Best arm in this test, greedy: turboderp's 2.00bpw pack plus the z-lab DFlash2 drafter, with a 2 GiB memory reserve. The box's twin GPU held other, idle models (5.5 GiB) the whole time. Window 24,737 tokens. Decode, one request (Ash's bench_concurrent): 145.5 tok/s code, 88.1 tok/s chat (median of 3 reps). Without the drafter the same pack keeps a 42,074-token window but decodes at 26.6 tok/s. Cold prefill 543 tok/s on a 2k prompt (median of 5 reps). Token-exact: drafts matched serial on every check. Needle found in a 15,805-token prompt (one run, 31.4 s). Honest limits: a ~25K window is small, one request at a time only (two or four lanes are refused outright), and the 2.50bpw pack plus drafter fits but only gets ~3K. It was just a test, if a 16GB card could run and what would be needed to get it to run.
4
7
913
I haven't even thought about that one. They will 100% see a price increase in the coming months. Apple markets them to the AI Crowd.
Welp I was on the fence but with the Spark price hike I’m in on a 256GB M5 Ultra Mac Studio! If you want that or a 128GB M5 Max MacBook Pro (or both! 😉) you should move now, the price will be going up soon. Also I expect the M5 Ultra 512GB just got a price bump.
1
519
I'm not sure if I will be able to own a house. But I will leave my kids with four Sparks and a 5090. These little shits better be grateful! 😂
128 GB DGX SPARKS ARE NOW RETAILED FOR $6,950 !!! 🤯🤯 What is going on ??????
1
6
245
4,999$ for the 64GB version. Everybody who bought a Spark in the last couple of months should be very happy right now! 64GB Versions and the new RTX Laptops will increase awareness for local AI. And everybody who is serious about it will soon realize that you probably want more than 64GB or the 128GB in the new RTX Spark models. Happy I got my 4 Sparks now. That would have been a lot more expensive today..
NVIDIA announces 64 GB DGX Sparks!! 😲 Starting Friday, Oct. 23 The new configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI. Priced at $4,999
3
4
898
Niklas retweeted
Silicon from sand qualifies. If the people stay quiet, these circuits will still testify. Luke 19:40 holds.
326
2,370
20,815
1,340,423
Fixed! spark3 and spark4 read prompts slower than spark1 and spark2. The cause was one stuck network card. I ran Mia's GLM-5.3-Flash recipe v1.3 (978b225, newer versions exist now) on ash's TensorFold 0.6.0, measured with her sparkDash and her 2,200 MHz cap. Prefill (how fast a box reads your prompt) on spark3 and spark4 was only 72 to 74 % of her README numbers. On spark1 and spark2 it was 98 to 99 %. (Prefill median of 2 runs, prompts of 8k to 128k tokens.) Cause: spark4's ConnectX-7 card got stuck after its cables were plugged in while the box was running. It sent only 13.2 Gb/s instead of 111.6 (ib_write_bw, spark4 to spark3). Not the cause: the IOMMU setting (spark3 has the same one and sends at full speed), core pinning, the GPU. Fix: one restart of spark4 with the cables plugged in. Now it sends 111.6 Gb/s and prefill is 100.4 to 101.7 % of her README. No outside load hit the boxes during the run. Decode did not get worse: 103 to 108 % of before (decode median of 3 runs; 1 to 4 requests added up, prose and structured). Limits: one pair of boxes, one fix. Why the card hangs, I don't know. If you plug ConnectX-7 cables into a running Spark: restart the box and check ib_write_bw in both directions. Mine was slow in one direction only.
3
2
540
If you are like me and you only read about 10% of your agents output. Read this!
This is 💯 Next time don't ask your agent for a report in markdown. Instead, ask for a beautiful well constructed HTML, a diagram, or even short video explainer. Oversight is becoming the main human job.
1
3
482