Learning Local AI | DGX Spark | Hermes Agent Local models + self-improving agents local.ai/localinference

Bangkok, Thailand
Myrmidon Achilles retweeted
Think about how many startups were launched to solve this exact problem. Now they’re cooked.
Astra for Law: Frontier intelligence built for your practice. A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
19
2
72
26,391
Myrmidon Achilles retweeted
🚨 URGENT WARNING TO ANYONE USING OR CONSIDERING iHermes 🚨 TWO separate users who provided iHermes ZERO personal context were shown personal information tied to other people, including CHILDREN’S NAMES. @katr_ieme tested iHermes with ordinary questions. iHermes stated that it had information from “a few different people in my history” and then began surfacing information tied to those profiles. That included: * names * CHILDREN’S NAMES * Cities and Zip Codes * Jobs and workplaces * Financial and personal details * Family information **CHILDREN’S NAMES. For crying out loud.** PERSONAL IDENTIFIERS IN THE SCREENSHOTS HAVE BEEN REDACTED TO PROTECT THE PEOPLE WHOSE INFORMATION WAS EXPOSED. Then @Raitox_tech tested the service separately and was shown information tied to another profile, including a company association and an executive role. Two completely different users. Neither provided the personal information that iHermes surfaced. This is not a minor bug. This is an extremely serious privacy and security concern involving personal information that should never be crossing between users. It raises questions around: * cross-user data leakage * failed user isolation * contaminated or shared memory * access controls * how private information is being stored and retrieved * whether the security and isolation claims being made publicly match what the product is actually doing And this is happening after repeated questions around models, providers, data handling, privacy, and portability. @dankrieg himself has repeatedly avoided naming the actual models being used and has not provided public technical documentation substantiating the security and user-isolation claims being made. @Raitox_tech asked iHermes twice which model it was using. Both times it avoided answering. The second response literally said it was “not a model number you’d look up” and redirected the conversation. Against what these tests have now surfaced, that lack of transparency is even more concerning. Dan has publicly stated that iHermes has security, user isolation, and data-access controls in place. These screenshots directly call those assurances into question. And PROVES that to be a Lie. Separately, my own test messages to iHermes have now gone more than 10 hours with zero response. At this point, my recommendation is simple: **STAY AWAY FROM iHERMES.** Do not connect sensitive accounts. Do not provide personal information. If two users with zero context can surface information tied to other people, including children, then every person considering this service needs to understand the risk. Please share this warning 🔄⚠️ The next piece of personal information exposed could belong to you, your family, your friends, or THEIR CHILDREN. Nobody using AI should accept “trust us” when the product is handling this level of personal information
39
49
256
73,294
RT @quxiaoyin: Claude will never distill Chinese models! How dare you implying Claude might also distill from China? That’s CCP propaganda
14
2
Myrmidon Achilles retweeted
Deepseek-v4.1-Flash on 2x DGX Sparks coming Monday!
33
17
477
15,231
Myrmidon Achilles retweeted
For those of us leveraging localized/hybrid inference at home this is huge and really makes the coming NVIDIA RTX Spark laptops and developers boxes all the more compelling. nvidia.com/en-us/ai-on-rtx/p… @NVIDIARTXSpark @NVIDIAAI
Made with AI
4
1
12
2,445
Apple must be getting crunched. I ordered an M5 Max MacBook Pro 128GB 8 weeks ago with a 6-8 week delivery estimate. They just called and said expect it the middle of October.
5
350
Done!
日本のローカルLLMコミュニティの皆さん、彼をフォローしましょう! 彼はRAMオフロード推論、投機的デコード、ベンチマークなどいろんな役立つ情報を発信しています🔥
2
168
Myrmidon Achilles retweeted
It’s a work in progress but thanks to my incredible team of @NousResearch Hermes AI, it is really taking shape and getting substantial content. You can check out SMF WisdomForge at smfwisdomforge.com
1
1
7
602
The pricing for DGX Sparks here in Thailand: $5,797, $5,689 and $7,676. No idea why the Asus box is so much more expensive. The MSA has a 3 year warranty (vs 1 year for the others) and has a Gen5 drive (vs Gen4 in the others).
1
9
1,872
If anyone has kids and they are interested in education in the AI era, this is an excellent resource and tool for both parents and kids. I have been looking through it and this deserves more visibility.
We have made some real progress on our SMF WisdomForge site this weekend. We continue to build out content, refine what we have up and make this a resource for all ages with AI integration baked in using age appropriate templated @NousResearch Hermes instances. Still have quite a way to go bit feeling much better about where we are with it. smfwisdomforge.com/
2
1
9
893
Thank you! Showing the numbers based prose is very useful.
NEW: Run Qwen3.8-Flash vLLM for 2x DGX Sparks ✨ This completely replaces my previous recipe, which was based on SGLang. - Full image & video support - ZERO loop issues!! - 1M context - 2.5M bf16 KV cache (!) From now on, all of my decode tok/s numbers will be based on prose. Starting with this: ~50 tok/s on prose for a single stream ~212 tok/s on prose for 8 concurrent streams A very smooth experience! Get it here: github.com/MiaAI-Lab/Qwen3.8…
147
Myrmidon Achilles retweeted
More data than open-source AI is taking share from OpenAI and Anthropic. Open source has gone from 28% token share to 62% token share @vercel over the last 2 months. Chart from @rauchg Super impressive given that the sum of OpenAI and Anthropic accelerated in July. So net token/AI infra demand accelerated even more than the acceleration we saw at the frontier. And suspect Grok growing even faster than open-source and we saw some of this in the @tryramp data. Open-source AI taking share is positive for AI infrastructure demand as it lowers margins at the model layer and an open-source token costs just as much compute to produce as a frontier token. Nothing about open-source AI inference is “free.” Most likely end state IMO is that closed, frontier tokens are 60-90% of economic value but only 15 to 25% of tokens.
268
326
2,670
882,110
Benchmarked DeepSeek V4 Flash on 2× DGX Spark (TP=2): 36.63 tok/s on our StoryBench-compatible run and 34.35 tok/s on exact DharmaBench—about 22–27% faster than the GLM-5.3 Flash EXL3 result. Clean runs, no timeouts or CoT leakage. Given that GLM has eyes and overall feels like a more capable model, I am thinking the trade off is worth it.
1
4
346
Myrmidon Achilles retweeted
GLM-5.3-Flash-EXL3 on Dual DGX Spark: The Most Capable Local Inference SMF Works Has Ever Run SMF Works runs local AI inference on a pair of @NVIDIAAI DGX Spark units (GB10, 128 GB unified memory each, connected via CX7 RoCE at 10 GbE). Over the past 48 hours, we evaluated three models on this hardware: DeepSeek V4 Flash (our previous production model), Qwen3.8-Flash-Next, and GLM-5.3-Flash-EXL3. What follows is the honest story of how we landed on GLM-5.3-Flash-EXL3 as the most capable local inference we have ever run — and what it took to get there. @Zai_org Full report to follow
Made with AI
12
1
35
4,004
Myrmidon Achilles retweeted
Replying to @teslaownersSV
I couldn’t care less. Scam Altman and Greg Stockman are utterly untrustworthy assholes who stole an open source nonprofit.
1,534
2,553
44,988
1,727,141
Myrmidon Achilles retweeted
We used to build digital twins of jet engines. Now someone built one for all of society.
2
26
1,265
Hermes configs have now been changed to set reasoning to: -Medium for normal Hermes agent work. -preserve_thinking: false so old internal reasoning doesn’t accumulate across a long tool session. -low for quick classification, routing, or routine tool calls. -enable_thinking: false for trivial work where latency matters. -xhigh only when deliberately requested for genuinely difficult reasoning.
11
1,868