Playing with my DGX Spark, BRB.

Middle of the GPU
BTW, I haven't updated the community lately regarding what I'm working on mainly because of some IRL setbacks but I did tried to simplify it as much as possible up to the stage that I build a second parallel app because the main one became to stuffed with features to be easy to manage by newbies. The new version is much more narrowed down to get actually great results without to many other variables.
There are so many awesome people in our local AI community building in the open, so wth, I guess it's my time to show you guys what I'm spending my tokens on lately and how I managed to keep my only DGX Spark burning in A/B testing 24/7 in the last 2 weeks. Many more screenshots to come so fallow the thread.
1
13
2,548
AgentSparko 💥 retweeted
First 10,000 users gets the lifetime access free with founders code (instruction on how to get it) on onlyprompts.ai
17
15
136
621,951
AgentSparko 💥 retweeted
r0b0tlab/GLM-5.3-Flash-EXL3-2.25bpw-sm121 + 3.00bpw DFlash2 drafter + native ExLlamaV3 runtime! Tested on a single @NVIDIAAI GB10 ♥️ At 8K C1, code: 22.61→49.96 tok/s with K=5 Exact single-key retrieval at 259,993 prompt tokens Extended eval suite in progress 🤓
5
7
32
2,131
AgentSparko 💥 retweeted
I’ll also be releasing the custom built app for training DFlash spec drafters, training models and abliterating models! Exclusive to paid Patreon subscribers for now. Getting closer to our goal of 500 paid subs to start abliterating bigger models! patreon.com/cw/AeonForge7
The Custom Aeon DFlash2 Drafter is coming along great! My retrained NVFP4 version is outpacing even the BF16 version now not just in Tok/s but even acceptance in many categories. Early release dropping for paid Patreon Subscribers then open wide release. patreon.com/cw/AeonForge7/me…
2
2
31
3,322
AgentSparko 💥 retweeted
Can a security defense catch an attack it hasn’t seen before? We teamed with @CrowdStrike to evaluate an offensive-defensive system built on its SafeMind agentic system, where AI agents simulate controlled attacks, turn telemetry into detection rules, then test them against new attack paths. Here’s how it works 🧵
38
88
608
71,841
AgentSparko 💥 retweeted
You were right, and this was the most useful feedback the Atlas got. It's live now: Every run opens with the answer: total tok/s vs speed per user, first-token wait, and how it was measured, concurrency, prompt/answer length, context window, speculative decoding, right at the top. 
Results are split by kind of test and sorted best-first. No more tok/s next to ms next to accuracy in one column. 
The engineering detail (flags, all metrics, raw record) is still there, but collapsed behind tabs. The recipe fits one screen on a small laptop. 
Every chart exports as PNG/SVG, and the home page is a living map. The 221 t/s Qwen3.8-27B run you mentioned: inference-atlas.0xbakeer.com… Keep the feedback coming.
Replying to @0xBakeer
Absolutely great inference atlas, but sir, you do need to improve the UX and make a more intuitive version of it as it is so confusing to browse through it and it really takes time to wrap your mind around it. I've seen the 221 t/s result in the main table for Qwen 27B and wanted to find the recipe and more basic info on it and first I had to scroll through all the 27B recipes that were not sorted by speed and in the same column there were different benchmarks measurements shown and in a total random order. I found the 221 t/s recipe and click on it to be again inundated with lots of info but I still could not see critical info like at what concurrency number, what context size those numbers were obtained and what was the total context pool. Then on the final recipe card (on a small laptop with browser tabs as a sidebar) 90% of the screen is filed with non useful information for 95% of the people and the box with the useful info shows just 5 lines of text and you have to scroll 28 pages of 5 text lines each to go through that info. All that info is amazing and thank you for your huge effort to build this and do all these benchmarks but the most important info for regular users should be on focus and the extra engineering info should be there just for those few that need it as secondary information because just a few users are engineers and will actually read through it.
4
3
9
986
AgentSparko 💥 retweeted
I’m building “The All Spark” a 36x DGX Spark cluster to run my own local agents and support the community with compute. 24 Sparks running today on a single cluster with the rest coming online after a quick power upgrade to the house 😅
219
57
1,053
151,640
AgentSparko 💥 retweeted
Replying to @mattshumer_
This is the type of discussions I see every few days in private groups. This guy comes here to say local models are useless and how good Opus 5.5 is while they nerf their models as soon as the new wave of subscribers cooled down. If this is a fair business model IDK what to tell you. This is just one of the reasons local models are crucial for anyone that is not an NPC.
1
7
595
Save this thread because it will be updated daily. The race is on for reverse engineering JEV from TypeSafe . ai There are multiple wonderful people working to bring on our local inference hardware a open source version of JEV. Someone even built this tracker that will help a lot to keep an eye on the new developments. I will also repost inside this thread my old post with other people attempts + cool stuff built on top of JEV. My only hope is @NousResearch pick up on it and does a deep integration loop in to Hermes that will allow us to boost a lot the speed at which our local agents work especially when we rely on local inference. huggingface.co/spaces/multim…
34
2
25
1,182
We have converted GLiNER2.5-Decide to coreml. ~4× faster, ~5× less Peak RAM, half the size. model: huggingface.co/FluidInferenc… code: github.com/FluidInference
1
2
51
AgentSparko 💥 retweeted
Less FOSS flame wars, more shipping Focus on what inspires you GB10 fam
2
5
298
AgentSparko 💥 retweeted
Qwen3.8-Flash-Next on a single @NVIDIAAI DGX Spark just got better 🚀 Default recipe 👇 ・Fixed a hidden bug that dropped per-layer embeddings: quality loss 1.397 → 1.344, same speed ・Follow-up replies 2x faster (1.66s → 0.81s), after tool calls 3.61s → 2.72s ・~49 tok/s prose, ~62 code (1 stream), up to ~168 tok/s (4 streams) ・Optional bit-exact decoding for evals and debugging ・Speculative depth 6 unlocked, +16% on code New opt-in lane on vLLM 0.30 (./start-v030.sh) 🌟 ・NVIDIA's official weights now fit on one Spark: the 48 GB per-layer table is read by the GPU straight from a file, built once and reused on every reboot ・FP8 KV backport: 801k tokens of context, 3/3 needles at 200k ・Best quality yet: NLL 1.332 ・Prefill +10% (2,140 tok/s), follow-ups 0.67s, code +6% at 4 streams ・Trade-off: prose ~25% slower. Use it for agents, keep the default for raw speed
14
10
80
30,152
Creating complex software with zero coding experience with Qwen 3.8 27B run locally on a DGX Spark using Hermes harness. The app is in production right now at the same time while I still develop it. When the app cannot be reloaded because of important on going tasks there is a dev tree that is modified and tested before updating the production instance. Qwen 27B decided on it's own to install "sensors" directly in our chat that notifies me and it of important events happening in the app. I have never seen this before and I did not ask it to do it but it uses that to do real time debugging without me even needing to prompt it that something bad happen. Actually, I don't even have time to see the problem and I see it already fixing it. I don't even do the development in Hermes TUI or Hermes Desktop, this is straight Hermes agent in my messenger app on my phone/computer. This stuff hits different.
Replying to @mr_r0b0t
Dude, when I checked my agent today after a power surge I found out that 27B installed sensors in our chat messages for a project it was working at. I could not believe that this came from the model itself and taught @NousResearch added it in one of the updates and ask it to investigate if that came from the model or from a Hermes skill or updates and it reported that it came from the model. No one asked it to do that and it did it just because I said something is important that was related to this. I'll make soon a thread about this as I evolved the idea already much further (think JEV like model + GPIO). Feel free to steal my idea ! Hope the guys from Hermes pick up on this message and they do it themselves too and if not I will publish something helpful for the community.
4
3
6
995
The projected API cost save info bubbles calculate using the last hour performance and the official Alibaba API pricing on OpenRouter.
1
74
Just to make an idea about the autonomy level of 27B developing in Hermes, I have 346 messages from the agent in the last 107 minutes. These are mostly tool calls and reasoning traces but also messages it sent me. This is why I like more the agent instead of the TUI because every decision it made it's documented and just one search away on my phone.
59
AgentSparko 💥 retweeted
Everyone is posting 3 and 4 Spark clusters this week. Here is what ONE DGX Spark does for one user, all measured on my box: MiniCPM5-2B, OpenBMB's drafter on: 100.8 tok/s Qwen3.6-35B-A3B: 89.7 Ling-3.0-flash: 69.5 Qwen3.8-Flash-Next: 55.5 DeepSeek-V4-Flash: 47.9 in EXL3 Gemma-4-E2B: 39.2 Qwen3.8-27B: 33.4 EXL3 + DFlash2 Tinfield-1 at 2-bit: 30.7 One model at a time, 256 tokens in, 256 out
22
10
79
4,423
AgentSparko 💥 retweeted
What if a small AI model didn't need to know everything, just how to think and where to look? That's Javi, a 4B model trained to think from first principles and reach for tools when it doesn't know. Here's how I built it, and what broke along the way 👇
Article

Building Javi: from pirate to first principles - a tiny model that reasons, checks, and never bluffs

After the pirate, I wanted a model that's actually useful. Inspired by @karpathy idea of teaching the models a "cognitive core", I set myself one goal: create a small, fast model that thinks from

20
3
32
4,844
You now have access to jev models that are better then jev in decisions! 🚀
49
Many people asked me about the GLi* evaluation tree, and this post is all about it.  Actually, the story starts before GLiNER. In 2022–2023, Chinese research teams and international collaborators explored several complementary approaches. UniMC (2022) framed zero-shot classification as multiple-choice prediction, encoding text and candidate labels together. USM (2023) matched schemas with text through token-level links, unifying entity, relation and event extraction. UniEX (2023) combined span detection, classification and relation extraction within one schema-driven extraction framework. On October 2023, Knowledgator published “As GPT4 but for token classification”, showing prompt-guided token classification for NER, question answering, relation extraction and other tasks. In November 2023, Urchade Zaratiana, and colleagues introduced GLiNER, using a bidirectional encoder to match text spans with entity labels supplied at inference time. Trained on the UniversalNER team’s Pile-NER dataset, it reported strong zero-shot NER results against the evaluated LLM baselines. I was excited about it, subsequently joined its development, and continue to co-maintain the project. In April 2024, Urchade and colleagues introduced GraphER, a related research direction that models entity and relation extraction jointly through graph structure. That June, Mykhailo Shtopko and I introduced GLiNER multi-task, extending the approach to question answering, summarization, and relation extraction. Our team also released GLiClass and bi/poly-encoder variants in 2024. In 2025, Jack Boylan and colleagues introduced GLiREL for zero-shot relation extraction; Robin Armingaud and Romaric Besançon developed GLiDRE. And Urchade and the Fastino team released GLiNER2, combining NER, classification, relation extraction, and structured extraction. In 2026, our team introduced GLiNER-Relex for joint entity and relation extraction. Fastino continued with GLiNER2.5 in August, adding boundary-based extraction, longer context, and constrained joint decoding. And this September, we released GLiFormer. It brings entities, relations, classification, PDF processing, and nested records into one framework with a shared encoder. For hierarchical extraction, it grounds values in the source, groups them into records, and predicts their relationships before assembling JSON. I’m proud of our contribution and grateful to all researchers, contributors, and users helping this ecosystem grow. Looking forward to more exciting contributions from the community.
16
updated Decision Index 0.1 → 0.2 🎯 better formula, +29 jev-like models, +21 benchmarks AutoJev-27B by @perplexity_ai CTO @denisyarats took the open lead, trailing jev by 0.8 points 🏆 come find the best model at every size, speed and use-case huggingface.co/spaces/multim…
7