Local LLM V100 96G & enjoy Omarchy; Tday unbug.github.io/tday/ ; CODELF (Github star 14k); MIHTool (Mentioned in Google I/O'13)

Los Angeles, CA
Pinned Tweet
Just a bit of insight: people are eager to build their own rigs for local LLMs, the same way we were obsessed with building gaming PCs 30 years ago. x.com/i/chat/group_join/g209…
1
8
4,936
Three PLX cards are already running in the PC—my 8× V100 mesh is halfway there.
1
13
763
On the consumer side, the Lisuan eXtreme LX7G100 already matches the RTX 4060 for gaming—at half the price. The MI AI Cube is planned for next year. Unfortunately neither supports CUDA.
what happens when China can manufacture AI hardware dramatically cheaper than the West?
2
1
4
813
The latest Unsloth Desktop allows to change LLM engine, looking forward to Strata in the list 😘
5
329
unbug retweeted
买不起英伟达 4090? 有个狠人只花了 200 多刀,直接在「二手矿机拆机件」上徒手焊出了跑大模型的方案。 这不是调参,不是用 Python 写脚本。 他是拿着电烙铁飞线、用底层 VHDL 硬件代码,把 Qwen(通义千问)直接硬编码写进了 FPGA 芯片里…… 算力丐帮的终极形态出现了,看完直起鸡皮疙瘩🧵👇
14
21
183
24,911
Damn, when I run Strata, my CPU turns into a jet engine—louder than my four V100s. Didn’t see that coming.😂
1
209
4 concurrency tested result. Still twice as fast as llama.cpp on my hardware.
Replying to @unbug
有没有测试并发能力?如果 多并发还能维持 80-90 速度,那是相当舒服了
1
1
396
Speed is always unbeatable—running Strata with DeepSeek Harness feels like DeepSeek V4.1 Flash itself. But it seems Max Effort isn't kicking in; the output doesn't really resemble Qwen3.8-Flash Q4_K_M_XL at Max Effort.
Finally got time to test Strata on multiple V100s. Answer to your questions: vs a single 32GB V100, it’s ~3× faster decode and ~3.5× faster prefill. Yeah, more VRAM wins. Caveat: this test isn’t fully fair—I’m still on Windows, and all 4 V100s share one PCIe 3.0 x4 link. I’ll rerun once my Linux build is ready later this week.
9
736
Finally got time to test Strata on multiple V100s. Answer to your questions: vs a single 32GB V100, it’s ~3× faster decode and ~3.5× faster prefill. Yeah, more VRAM wins. Caveat: this test isn’t fully fair—I’m still on Windows, and all 4 V100s share one PCIe 3.0 x4 link. I’ll rerun once my Linux build is ready later this week.
6
3
20
2,676
So far, so good — both MoE and dense models passed. No crashes, stable temps, clean output. The cheap ECC Error V100 is holding up better than expected...for now...for now.🥳
Sorry for the wait, everyone — the ECC-error V100 results are in. Both cards were detected, drivers installed fine, and temps were stable. GPU2’s ECC errors keep climbing: totally dead. GPU3 only has static errors at boot: stable and usable. I ran Qwen3.8-Flash Q4_K_XL on GPU3 with Unsloth + llama.cpp and did a full test on both cards. Report’s in the image. One good news, one bad news — but the total was still under 30% of one healthy card. Honestly? Pretty happy. 😅
2
343
The 2× V100 32GB cards with ECC errors are here. Time to put them in and see what happens. Half excited, half scared—heart pounding like I’m meeting my crush. I’ve said my prayers 😅
Just ordered 2× V100 32GB with ECC errors. Damn, they're 6x cheaper than normal ones. Some says ECC errors won't matter much for local LLM. Anyone have experience with this? Wish me luck 😅
8
4
51
4,996
Sorry for the wait, everyone — the ECC-error V100 results are in. Both cards were detected, drivers installed fine, and temps were stable. GPU2’s ECC errors keep climbing: totally dead. GPU3 only has static errors at boot: stable and usable. I ran Qwen3.8-Flash Q4_K_XL on GPU3 with Unsloth + llama.cpp and did a full test on both cards. Report’s in the image. One good news, one bad news — but the total was still under 30% of one healthy card. Honestly? Pretty happy. 😅
2
180
Sorry for the wait, everyone — the ECC-error V100 results are in. Both cards were detected, drivers installed fine, and temps were stable. GPU2’s ECC errors keep climbing: totally dead. GPU3 only has static errors at boot: stable and usable. I ran Qwen3.8-Flash Q4_K_XL on GPU3 with Unsloth + llama.cpp and did a full test on both cards. Report’s in the image. One good news, one bad news — but the total was still under 30% of one healthy card. Honestly? Pretty happy. 😅
The 2× V100 32GB cards with ECC errors are here. Time to put them in and see what happens. Half excited, half scared—heart pounding like I’m meeting my crush. I’ve said my prayers 😅
3
3
35
2,163
So far so good, running Qwen3.8-flash on them right now, and will run a full ECC ERROR testing with Qwen.
2
232
When you're tinkering with an open-frame rig in your bedroom, motherboard standoffs are an absolute must—you'll be tearing it apart and rebuilding it over and over until the layout finally feels right.
The perfect case for my V100 mesh: large enough for 4 pairs of NVLink V100s, small enough to fit under my desk.
2
9
954
unbug retweeted
Please share this to support Strata development 👇 Looking for sponsors or rental of at least: - PC with AMD GPU + 64GB RAM - PC with Intel GPU + 64GB RAM - PC with Dual GPU setup + 32GB RAM Location: Slovenia (Europe) I would like to work faster as we’re only 30% in. Please message me Thank you Niko
32
93
406
27,084
Living in China, this is exactly how it is. No matter what you buy—even DIY parts—the quality is so overkill that you don’t even worry about it. You literally only ask: is the price right, do I like the design, and can they ship today?
therapist (opus): you can't just put 8 gpus in one computer, you should give up China:
4
395
I have not find yet due to the V100 is a data-center GPU. But you can easily build one with a dual-tower X99 case, and there are even cases made specifically for V100 DIY builds.
Replying to @unbug
do you know if there is a workstation that can support 4 v100 out of the box? preferably with another slot for a low profile gpu for display?
5
1
25
2,755
These cases are made for dual E-ATX motherboards.
160
I have not posted these cases due to I don’t have room for them in my bedroom 😂
1
232
In the last 24 months, I spent 5x more on online SOTA than on my entire local LLM build. Had I known open-weight models would get this good, I'd have bought a 5090 and a DGX Spark before wasting all that money.
Replying to @unbug
A fifteen-hundred-dollar box running the current best open model is quite something. I'm API-only now and even I paused on that.
4
5
514