Learning Local AI | DGX Spark | Hermes Agent Local models + self-improving agents local.ai/localinference

Bangkok, Thailand
Based in Thailand
Filter
Exclude
Time range
-
Minimum likes
Replying to @MichaelGannotti
You know, the last time you were on holiday models rained down on us 🙏😀
1
3
32
Apple must be getting crunched. I ordered an M5 Max MacBook Pro 128GB 8 weeks ago with a 6-8 week delivery estimate. They just called and said expect it the middle of October.
5
351
Done!
日本のローカルLLMコミュニティの皆さん、彼をフォローしましょう! 彼はRAMオフロード推論、投機的デコード、ベンチマークなどいろんな役立つ情報を発信しています🔥
2
170
Replying to @NeoAIForecast
This has pretty much always been the case. Lot of love for OC since it put me on this path but reliability has always been a problem.
1
2
111
The pricing for DGX Sparks here in Thailand: $5,797, $5,689 and $7,676. No idea why the Asus box is so much more expensive. The MSA has a 3 year warranty (vs 1 year for the others) and has a Gen5 drive (vs Gen4 in the others).
1
9
1,872
Replying to @runsonai
It is $7,676 in Thailand right now...up for $5,500'ish about 3 weeks ago.
1
2
1,040
If anyone has kids and they are interested in education in the AI era, this is an excellent resource and tool for both parents and kids. I have been looking through it and this deserves more visibility.
We have made some real progress on our SMF WisdomForge site this weekend. We continue to build out content, refine what we have up and make this a resource for all ages with AI integration baked in using age appropriate templated @NousResearch Hermes instances. Still have quite a way to go bit feeling much better about where we are with it. smfwisdomforge.com/
2
1
9
893
Replying to @runsonai
I agree, I have 4 agents in non-stop training runs for more than 36 hours straight using DSV4F and it is rock solid.
1
4
788
Replying to @jun_song
I am only sticking with it for full computer use.
2
679
Benchmarked DeepSeek V4 Flash on 2× DGX Spark (TP=2): 36.63 tok/s on our StoryBench-compatible run and 34.35 tok/s on exact DharmaBench—about 22–27% faster than the GLM-5.3 Flash EXL3 result. Clean runs, no timeouts or CoT leakage. Given that GLM has eyes and overall feels like a more capable model, I am thinking the trade off is worth it.
1
4
346
Replying to @CardilloSamuel
I really don't view it as a high percentage play to piss of the richest human on the planet and also, potentially, the smartest.
2
55
Replying to @jun_song
Open source has to win.
1
2
332
Replying to @twid
Perfect break for dinner in there 👍
2
43
Hermes configs have now been changed to set reasoning to: -Medium for normal Hermes agent work. -preserve_thinking: false so old internal reasoning doesn’t accumulate across a long tool session. -low for quick classification, routing, or routine tool calls. -enable_thinking: false for trivial work where latency matters. -xhigh only when deliberately requested for genuinely difficult reasoning.
11
1,868
New finding (to me at least)..Qwen3.8 in Hermes defaults to xhigh reasoning unless Hermes spefically specifies a different reasoning level. In Hermes at least, xhigh will make it think indefinitely without generating an answer...well tested up to 4.5 hours and 4M tokens burned purely thinking.
4
588
An update in using Hermes for everyday business tasks: We are in the 3rd round of tuning a Hermes bot to specialize in building financial models. Currently, there is an orchestrator bot overseeing an analyst bot, both running DSV4F. After analyzing the output with Codex, we are moving the analyst bot to Qwen 3.8 27B to compare error catching efficiency. I am guessing we will also wind up splitting the workload between 1 Hermes bot that is tuned to create assumptions pages, and another Hermes bot that is tuned to produce the model.
8
14
1,916
Replying to @MiaAI_lab
I think the Spark is still better with higher concurrency use cases?
5
843
Replying to @mweinbach
The Studio Ultra with 256GB and 1TB is $10,799. Not as bad as I was expecting.
1
14
4,625
Replying to @TechMDAI
I may breakdown and buy a non-mac machine for the first time in 10 years to try it.
2
28