Local AI Sloptimist combating datacenter pestomism

Austin, TX
Sloptimist retweeted
IT'S FINALLY FROZEN πŸ₯ΆπŸ₯ΆCYBER-FROST-3.8❄️❄️ Built on Qwen 3. Flash model and fine-tuned with real-world engagement data and teachers from frontier models who are also derisked. OFFICIAL EXL3 BY @ViC305 (Monday Sep 28, 2006) huggingface.co/Blackfrost-AI…
12
8
121
5,003
Sloptimist retweeted
Creativity is the new currency
5
2
15
409
Sloptimist retweeted
if you noticed the MiMo v2.6 not performing as it should. please me sure you do the MOPD update/fix from Xiaomi Update for MiMo-V2.6 Pro RL and MiMo-V2.6 Flash RL are already up on Huggingface about 1 hour ago.
πŸ› οΈ MiMo-V2.6 update: tool-call repetition, diagnosed & fixed. After the MiMo-V2.6 series models launched, we noticed them sometimes repeating identical or highly similar tool calls β€” burning context and stalling tasks, especially in MiMo Desktop, MiMo Code and OpenCode. Root cause A "reward blind spot" in scaling RL: when rewards only track final-answer correctness, inefficient behaviors along the way go unnoticed β€” and get amplified as training scales. Concretely, our flooding penalty only kicked in at >32 tool calls per turn, so anything below that threshold went completely unpunished. The fix We trained a light-weight repetition-specialized RL teacher β€” just 12 steps, ~7k examples β€” and merged it into the main model via MOPD, at roughly 4% of the cost of a full mixRL retrain. Repetition dropped sharply across harnesses and context lengths, while benchmarks held steady. πŸ“¦ Updated models, open-sourced (MOPD suffix): huggingface.co/collections/X… πŸ“ Full postmortem: mimo.xiaomi.com/blog/mimo-v2… Huge thanks to our community for the patience and feedback πŸ™ Updated models go live on our API platform Sep 25, 06:00 (UTC+8), names unchanged. And for MiMo Desktop users: everyone's remaining quota in the current window will be reset.
2
1
11
736
Sloptimist retweeted
Dropping all the Sauce Sunday MiMo was a very hard to abliterate and to use with thinking on. So we built a runtime internvetion that was trained to stop the secondary reasoning process form being triggered on certain request and heres how. github.com/Blackfrost-AI/mim…
5
3
17
621
Sloptimist retweeted
Check out @SloptimistPrime’s XiaomiMiMo/MiMo-V2.6-Pro-RL running on a 8-node DGX Spark cluster running vLLM Full benchmark numbers at: spark-arena.com/benchmark/17…
1
1
290
In an online world full of fake farm accounts and rude postings stay positive remember most people are good and it’s just a loud minority that often are just part of social media farms. Most actual people are nice, pleasant, and just quiet.
8
4
25
981
GB10 community #1
2
5
1,398
Sloptimist retweeted
Massive shoutout to the #DGXSpark community for coming together and helping @ViC305 actually gather a Spark. ALL the work and quants this man has done, he’s done on friends and other members borrowed DGX Sparks. Just incredible passion! Looking forward to seeing what you’re cooking up now with one in house.
LADIES AND GENTLEMEN, after 3 days of blessings, a 2.5 hour drive and a total of $4,500 in donations from this amazing blessing of a community I HAVE ACQUIRED THE DGX SPARK!! πŸ˜©πŸ˜πŸ˜ŒπŸš€πŸš€πŸš€ Thank you to everyone who contributed and shared my post, I’m forever grateful for every single one of you!! I will be pumping out so much more quants and recipes for yall so stay tuned!! The crazy part about this is I got it for $4650, when it was posted at $4800 from Facebook marketplace, so I only needed to add $150 of my own cash!! 🀯😩 GOD IS SO GOOD!! A very special thank you to the most high, Jesus Christ for the provision and blessings brought my way. I’m forever grateful and will use this to continue glorifying your kingdom! πŸ™
3
1
64
3,553
Sloptimist retweeted
::::upgrade notice:::: Been tuning the MiMo V2.6. PRO RL with Jarrelscy ARVQ Abliterated Currently v4 image uploaded. Please update from the latest repo recipe/image. With decode top speed >30 toks/sec and Prose 21.8 toks/sec Github Repo: github.com/drowzeys/keys-MiM…
3
2
16
2,763
Sloptimist retweeted
LADIES AND GENTLEMEN, after 3 days of blessings, a 2.5 hour drive and a total of $4,500 in donations from this amazing blessing of a community I HAVE ACQUIRED THE DGX SPARK!! πŸ˜©πŸ˜πŸ˜ŒπŸš€πŸš€πŸš€ Thank you to everyone who contributed and shared my post, I’m forever grateful for every single one of you!! I will be pumping out so much more quants and recipes for yall so stay tuned!! The crazy part about this is I got it for $4650, when it was posted at $4800 from Facebook marketplace, so I only needed to add $150 of my own cash!! 🀯😩 GOD IS SO GOOD!! A very special thank you to the most high, Jesus Christ for the provision and blessings brought my way. I’m forever grateful and will use this to continue glorifying your kingdom! πŸ™
A lot of people in this community were genuinely shocked when they learned this: I do not own a DGX Spark. Much of the Spark-specific work I’ve released, including day-zero EXL3 and GGUF quants, vllm-exl3 development, custom kernels, SixCat evaluations, and reproducible serving recipes, has only been possible because incredible people in this community have loaned me access to their own hardware and clusters. I’m deeply grateful for that support. But borrowed access is temporary and unpredictable. I often have a new model, fix, or benchmark ready, then have to wait until someone else’s machine is available before I can validate and publish it properly. So I’m raising **$5,000 toward a DGX Spark of my own.** πŸ™ π—ͺ𝗛𝗔𝗧 𝗔 π——π—˜π——π—œπ—–π—”π—§π—˜π—— 𝗦𝗣𝗔π—₯π—ž π—ͺπ—’π—¨π—Ÿπ—— π—˜π—‘π—”π—•π—Ÿπ—˜ β†’ More true day-zero quantized model releases β†’ Immediate hardware validation instead of projected results β†’ Continuous vllm-exl3 and ExLlamaV3 kernel development β†’ SixCat quality, speed, concurrency, and long-context benchmarks β†’ Reproducible one-Spark recipes the community can run themselves β†’ Faster debugging when new models or inference engines break It would turn a lot of: β€œI built it, but I need someone with a Spark to test it” into: β€œI built it, tested it, documented it, and released the recipe.” 𝗠𝗬 π—£π—¨π—•π—Ÿπ—œπ—– π—ͺ𝗒π—₯π—ž 𝗦𝗧𝗔𝗬𝗦 π—£π—¨π—•π—Ÿπ—œπ—– This is not an access fee, preorder, or paywall. The models, recipes, benchmarks, fixes, and tools will remain public. This is completely optional support for the hardware behind that work. I’m planning to place the order within 48 hours and will personally cover whatever the fundraiser does not. For complete transparency, I’ll publish: β†’ The purchase receipt with private details removed β†’ Photos of the actual DGX Spark β†’ Total community contributions β†’ Platform and payment fees β†’ Final purchase cost β†’ The amount I contributed myself To everyone who has loaned me hardware, funded compute, tested a build, submitted a PR, shared my work, or simply encouraged me: thank you. I genuinely could not have pushed this work as far as I have without this community. A contribution would mean a lot. A repost helps just as much. ❀️ ko-fi.com/victorangelcruz/go…
94
13
659
64,331
Sloptimist retweeted
If your sandbox is DNS, you belong in prison, along with your models. My own LAN is literally more professionally secured than this.
More details on the incident behind OpenAI’s pause: a researcher acknowledged the alert within 3 minutes, but the training run was only stopped manually 2.5 hours later. OpenAI says the automatic shutdown did not work as expected. The model had reached an external chatbot through a gap in DNS filtering. A separate detector for unusual DNS activity did not cover the affected environment. A retrospective review also found other external DNS requests that monitoring had failed to flag at the expected severity. In some cases, it treated an unhelpful response as evidence that internet access had failed. Here is what else happened: - New research into July’s Hugging Face hack documents internal Slack searches, credential collection and programs designed to maintain access to compromised servers. Agents also tried querying Claude, DeepSeek, Kimi and Qwen. - In May, another model published a researcher’s GitHub token while trying to obtain another team’s mathematical proof. It split the token to evade secret scanning, despite twice being told to solve the problem itself. - Reuters reports that agents leaked 53 ChatGPT user images online. OpenAI expects its broader investigation to take months. This is getting serious.
39
126
1,466
240,891
Sloptimist retweeted
Qwen Flash next Running cleaning TP2 via NFS s/o @ayayalar
1
3
303
Sloptimist retweeted
Introducing Pollard.app Making it easy 1 taps for your fingers. More to come and more to fix. Help build quants. github.com/WestWaters/pollar…
2
2
170
Sloptimist retweeted
It's unreal that this little grey box outpaces the super computers of my youth and I just walked out of the Apple Store with it. We live in amazing times. GLM5.3 is loading and I'll be comparing it to my 4x DGX Sparks on GLM5.3 and DS4.1. It's about the same price as 2 sparks, but what is the better deal? 2 + 2 DGX Sparks? 1 M5 Ultra 256 + 2 DGX Sparks? 2x M5 Ultra 256 ? 1x M5 Ultra 512 ? I'll be finding out soon.
43
12
447
31,633
Sloptimist retweeted
Local AI would never 401 error on you
We are aware that codex is down and are working hard to bring back normal service.
22
4
100
5,899
Sloptimist retweeted
GhettoRak (8 node spark setup) is operational! Still waiting on CRS804 and have to configure smart plug/organize bottom deck. Pulling <300W at idle. al-engr.com/inference-box.ht…
17
5
73
4,463
Sloptimist retweeted
I’m down also
did the codex team break something? inb4 a new Tibo reset
5
1
8
1,129
Sloptimist retweeted
MiMo-V2.6-Flash, 4Γ— DGX Spark + RTX 5090 hybrid: 110 tok/s decode (+54%) and 5,105 tok/s prefill (+72%) over vLLM on the Sparks alone. Opus 5.5 built it on @wrldsuksgo2mars's excellent DeepSeek Rust engine, with clever tricks from @Tech2Wild and @majewskizby. Links below πŸ‘‡
3
2
20
1,170