Advance. By any means necessary.

What features are we missing from TensorFold?
51
4
57
8,530
Faster speed. It will end everything.
2
2
361
Take speed to its absolute limit—vLLM and SGLang won’t even be fit to tie your shoelaces!
20
The line between filmed and rendered is getting blurry. This mosaic doesn't exist. It's 127,445 tiles, each its own 3D object, placed one by one in code by Claude Code with Opus 5.5. Starts inside her eye, ends on the pearl!
7
7
246
18,526
What is the prompt? It’s so beautiful!
1
1
176
This aged well. You could’ve bought a DGX Spark for $3500-4000, and Codex's $200 plan was actually pretty generous. 3 month later we have a 64 GB DGX Spark retailing for $4950, the 128 GB retailing for $6950, and a new $500 Codex plan. How will it look like in 1 year from now?
"It just doesn’t math." I keep seeing this take. But why run AI on your own hardware? Because for many of us, the cloud just doesn’t math either. Personally, I replaced all my Claude + Codex usage with DeepSeek-v4-Flash running locally on two DGX Sparks. For reference: $8,000 spent on the Claude Sonnet API would currently buy you roughly 800 million output tokens — about 5 months of continuous generation at 60 tokens/sec. The "It just doesn't math" claim: "For $8000 + electricity you could get over 4 years of Claude Max $200/mo sub plans, which would give you more Sonnet usage than your local setup." Who says Claude Max stays at $200/mo? They’re literally losing money on every sub right now — this price won’t last. They also nerf the limits constantly, & you’re STILL rate-limited even on the top tier. I know you can’t run actual Sonnet locally. That’s not the point. The point is: for my workflows, the local models I can run are good enough to fully replace it. A lot of people (including me) simply prefer not to send work through certain cloud providers, whether for privacy, trust, or other reasons. So the real question is: if the hardware can replace what you’re already paying for in API costs, is $8k justifiable? For me, it absolutely is. Plus… you actually own it. You don’t like owning things?
31
21
229
34,792
感谢dsv4f0731
1
115
Getting into local AI image generation! Starting with two recent open models: Ming-Image-0.1-Design and Qwen Image 2.1. This is the first post in a series of comparisons I'm going to do between them. I ran both on my RTX 5090. Same prompt, same seed. Both did a great job. Ming's image looks more cinematic. Qwen's is clean and simple. Two different styles, both usable imo. More tests coming in the next few days. Model links and full prompt below.
28
4
115
9,911
The former looks like only 720p, and the image quality feels very similar to minimaxh3. This round, qwen seems to have won.
141
For reference, this is what I paid for my 4 TB 128 GB DGX Spark in July. Madness!!!
NVIDIA announces 64 GB DGX Sparks!! 😲 Starting Friday, Oct. 23 The new configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI. Priced at $4,999
81
19
643
51,659
Nvidia saw this 128GB machine lagging way behind the 5090 in price hikes and decided to step in personally to fix it. Truly touching😭
1
173
We are going to release a new method to free up 2GB+ of RAM on DGX sparks, stay tuned. This is going to be a great win for TP=2 setups.
2
1
28
1,700
期待释放 20GB+ 的 RAM😚
12
Interact with the article here: cerebras.ai/blog/disaggregat…
5
8
4,267
@LisaSu 你们应该收购这家公司而不是骗子李飞飞 1500tok/s>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>高斯泼溅
128
Replying to @ItsmeAjayKV
希望越大失望越大 兄弟你不如期望它会排倒数第一🤣
25
Unless something new comes out in the next few days. these are all coming soon with TensorFold: Solo DGX Spark: - Qwen3.8-27B 2x DGX Sparks: - Qwen3.8 Flash Next 3x DGX Sparks: - GLM 5.3 Flash EXL3 - DeepSeek v4.1 Flash 4x DGX Sparks: - GLM 5.3 Flash EXL3 - DeepSeek v4.1 Flash
71
15
458
27,765
GLM5.3 Flash is nearly twice as fast on the new TensorFold engine. This is the funeral for vLLM and SGLang. Congratulations
1
2
363
GLM5.3 Flash is nearly twice as fast on the new TensorFold engine. This is the funeral for vLLM and SGLang. Congratulations @MiaAI_lab @ashxhart At this speed, it has almost caught up with Grok's speed and the speed of running small models on a 5070 Ti with 900GB bandwidth. You are the new king of local models.
1
41
Thinking humans can control a superintelligence is pure cope. A lesser mind cannot manage a greater mind. Simple as that.
JUST IN: Trump declares “whoever wins S.I.” — Super Intelligence — will rule the world.
89
10
199
11,172
也许他只是想永生
22
GLM-5.3 has helped defend 389 open-source projects, with 4,249 potential vulnerabilities found so far. OpenVuln is still running. The service remains free, and findings go privately to maintainers. huggingface.co/spaces/zai-or…
146
577
5,250
316,363
A畜越是反对,就说明你们做的越对。 加油,期待你们开源opus5.5,让这家反人类的公司倒闭
1
251
DeepSeek 弹性计算团队大量 HC 招人!尤其需要资深工程师。 来看看这篇技术分享:《DeepSeek 弹性计算 (DSec):面向大规模 Agent 训练的沙盒基础设施》 zhuanlan.zhihu.com/p/2088265…
59
55
694
105,698
A畜还没掌握的技术你们不要开源啊,不然全部偷走了还用计算资源碾压你们。 A畜已经会了的可以全部公开,opus5.5免费送这傻逼公司第二天就要倒闭
2
787
IQuest-Q1 lands with 320B MoE scale, 15B active parameters, and a 512K context window for long-horizon coding agents. 🤖 modelscope.ai/models/IQuestL… 🧠 Built for agentic coding, complex reasoning, and multi-step tool use across long-running workflows. 🏆 Scores 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo. ⚡ 256 experts with 8 activated per token, hybrid sliding-window/full attention, and recursive MTP for faster decoding. 🛠️ Production deployment supports SGLang and vLLM, with integrations for Claude Code, Codex CLI, and OpenAI-compatible APIs. 📜 Weights released under the IQuest-Q1 License.
7
13
98
6,594
得分还行就正常比,得分太难看就跟falsh比 跟ds4.0上一代模型比,这种不诚实的厂家的模型有必要下载吗?
1
142
Never in a million years would I have even dreamed my project would compete with the absolute best engines on the planet 🤯 Waking up to this is surreal 🫶🏼
Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥 This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements! - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared to vLLM! Performance: Decode prose 62+ tok/s single stream Decode prose 119+ tok/s on 5 streams Prefill is mostly 2500 tok/s across the board! In addition, expect TensorFold recipes for GLM 5.3 Flash and DeepSeek v4.1 Flash - coming soon! Get it here: github.com/MiaAI-Lab/Qwen3.8…
25
11
320
23,799
你,去干掉VLLM 振作起来,兄弟👏
1
144
paid shill btw
71
21
1,375
35,330
@thsottiaux 你就钱花这些玩意儿身上?它对的起你的推广吗?发的什么玩意儿
74
Something that bugs me about speculative decoding, it's supposed to be free speed, but it can change your answer. MTP lets the model guess a few words ahead and check them all in one pass. But checking several words at once rounds numbers slightly differently, so sometimes you can get a different word. On the M5 Ultra that's fixed now. Every checked word runs exactly the math it would run on its own, so MTP on and MTP off give byte-identical output, and MTP is still ~1.6x faster. Merged in oMLX.
6
19
1,539
或许可以设置100%一致再放行?
1
1
70
今天提名几个广泛使用、内容丰富、更新及时的 DSH 插件导航站吧~ dshfind.com/ dshmarket.com/ awesome-dsh-plugin.com/ 其中 dshmarket 也可直接作为 DSH 插件安装到设置页中。 (不代表公司立场,不对非官方网站的内容负责。)
从 DeepSeek 官方 API 处统计的数据来看,约有 60% 的 DeepSeek Harness 用户使用了至少一个第三方插件。第三方插件是 DeepSeek Harness 用户体验中最具特色且不可缺少的一部分。DeepSeek Harness 团队将持续支持第三方插件生态的繁荣发展,并推动插件 API 趋于稳定,在将来减少和尽量避免破坏性更新。 接下来的几天我个人将每天推荐一个优质的 DSH 第三方插件,欢迎 DSH 插件作者在本 thread 下自荐。我会结合插件质量及后台实际统计到的插件使用量择优推荐。 DeepSeek Harness 团队祝大家中秋快乐阖家幸福! (注:在用户使用官方 API 及模型时,DSH 会向官方 API 上报实际使用的插件包名和版本。此类上报不额外消耗 tokens。)
35
21
296
69,045
模型可以开源,技术方面还是藏点东西吧,A畜全偷过去了咋玩呢
1
364
We need one more Qwen 35B @QwenDevs 🙏🏻 Qwen3.6-35B-A3B unsloth Q4 is the first model on my 3060 that blew my mind. What a great model, it was soo good, MoE arch meant i can run with RAM offloading (experts), i even ran Q6 one on my 3060 + 64GB ram. Everyone is asking for it, many are giving up hopes, but you can give us one more right?
what was the first model you ran locally that made you think this is actually real?
13
13
166
10,723
I'm doing everything I can to get this released. They're actively testing and training an A3B MoE model checkpoint. glad you and others are keeping up the pressure publicly. genuinely helps! i'm hopeful.
13
8
204
36,383
兄弟,我相信你,但是你之前是不是才骗了我们一次?Anthropic研究员?😢
582