Personal AI. Local-first. For the inner life. 2× DGX Spark · M4 Max 128GB

Based in United States
Filter
Exclude
Time range
-
Minimum likes
LotusDecoder retweeted
说一个我的猜想吧, Opus5.5之所以忽然会制作视频或者完成度高的项目,并不是智力有很大的提升,而是他学会了。而且,他的老师就是各位消费级plan的用户。 先不说隐私协议,光是用户就是A社挑选过的。 更好的用户带来更好的学习素材,你当Claude是生产力,他在贪婪地蚕食你的知识,经验,品味,癖好
1
1
1
103
LotusDecoder retweeted
Replying to @Philo2022
我有个暴论,现在用token,等于12年买比特币,每个月都能把200刀套餐干光的人,未来几年不会穷到哪去的,不管token是用来学习了,还是用来build。
7
604
LotusDecoder retweeted
The TensorFold engine lands on the Sparks 🫡 Showing roughly 2x gain across the board with headroom for more. Point your agent at the repo, enjoy the speed 🚀🚀🚀 tensorfold.dev @NVIDIAAI @NaderLikeLadder @sundeep
TensorFold Inference Engine is here 🚀 I spent six months making one weight read count for more than one token on Apple Silicon. Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes. Qwen 3.8 27B MLX 4Bit - 120-124tks Nemotron Lightning MLX 4Bit - 188-206tks Qwen3.8 Flash Next MLX 4Bit - 88-92tks CUDA Implementation is in Alpha showing strong gains. The Repo is in the comments 👇🏼
14
11
138
8,354
Replying to @ivanalog_com
我不知道。
1
13
Replying to @ivanalog_com
之前尝试了一点,对 DGX spark GB10,跑了一天,没效果。 这几天是针对 软件工厂/软件流水线 ,在做 RSI 了, 找到了感觉, 原来 claude 的 6月报告诚不我欺, 一个环节接着一个模块,重构优化-测试-效果好-上线。 这样捋一遍,再总结成模板, 让 agent 去自动捋。 节省token 提高速度,提高通过率。 飞一般的推背感。
1
37
Replying to @ivanalog_com
对大部分人来说,这倒也是的, 我开始进入 RSI 阶段了, 现在连opencode go 的 deepseek-V4.1-flash 都要用干用尽了。
2
1
40
Replying to @ivanalog_com
利用率的点位在 RSI 时代,也马上没了。 有多少 token + cpu + 内存 + 硬盘, 转化为提升多少效率, 效率又有几率转化为效益。
1
3
120
哇,x 还是太厉害了。 可以直接对话塔勒布老师, 我还记得好几年前一边读黑天鹅三部曲, 一边看到美股有位50美分哥持续数月买地板价vix看涨期权, 直到某次极端行情,他压中了。 真是书本在现实的一次典型印证。
12
9
1,100
Replying to @realWeZZard
是 2080ti 22 g ?
1
382
LotusDecoder retweeted
jev and multi-agent swarm, it's all about scaling
13
584
Replying to @dongxi_nlp
恭喜😇🥳
1
174
发现现在agent编程, 带着一种 hack 精神, 如果提示词里的护栏不够多的, 一不留神, agent会用一种思路清奇和诡异的方式来实现, 然后也不能说是错的, 按提示词和测试来说, 全是合格的, 但是埋下了一个坑,很久才会发觉,原来当时是这么处理的。 🤪 所以有人说agent是被污染的黑圣杯,很形象。
2
7
757
Replying to @yifanxu_ephai
是的,人不知道,agent 为了解决一个点,悄悄干了什么, 即使是 端到端 测试, 为了搞定目标, agent会选一种非常神奇的路径来实现,我想都想不到。
3
89
耳机音响有煲机, claude账号也可以煲? 😅
Claude 避免封号焚决: 新号煞有介事地开始构建或者探讨极为艰深的科学问题。 Claude 后端必有一个账户打标机制。 它也想要科学家和高级工程人士的生产协作数据。 被封号,因为它判断你是一个麻瓜。 (本推文不改变 A 畜本质。咱们一起玩儿它)
2
4
1,293
LotusDecoder retweeted
The need for INSANE speed? Yes ⚡️ DeepSeep v4.1 Flash for 4x DGX Sparks is getting... ridicules 🤯 Decode improvements - Prose decode is now 87 tok/s single stream - On 4 streams it's 163.6 tok/s - Code decode is 124.8 tok/s single stream - On 4 streams it's 246.8 tok/s Prefill numbers? 5800-5925 tok/s on 16k-128k 5394 tok/s on 256k It feels like a FAST API! If you have 4x DGX Sparks, you need to try this out ASAP. I wonder if the M5 Ultra 512 GB could beat this. The improvements for now are only for TP=4 and do not effect TP=3. Thanks to @majewskizby for this awesome PR! Get it here: github.com/MiaAI-Lab/DeepSee…
32
18
275
23,164
我觉得很快会超级缺乏 cpu 硬盘 内存。 这波 国产三家 flash 的集中爆发。deepseek glm qwen。 速度更快,成本更低。 gpt luna 的量大管饱。 opus-5.5 的提质降价。 那么以前是任务能做完就ok啦, 现在很快要求,多快好省, 于是送进RSI打磨细节, 要开上并行沙箱, 所以大量需求cpu 内存 硬盘。
19
1
33
2,925
Replying to @ScarletKc
不知道啊
1
377
Replying to @gantrols
🤣
373
opus-5.5 的缺点也是有的, 主打一个容易丢三落四, 某些细节忽略掉了, 我是发现两回了。😅
15
29
5,103