Building with AI & LLMs | Insights, recipes, tools & honest experiments

Pinned Tweet
Come check out my website. It includes a list of my latest & greatest recipes. The "Install Model" button gives you a ready-to-use prompt for Claude, Codex, or others. It adapts the recipe to your setup, starts the server and tests it. mia-ai.net/models
109
130
1,281
192,848
Opus 5.5 with Ultracode (had to restart) Currently running 32 agents, with 9.2M tokens burnt.
First time running Opus 5.5 with Ultracode Let's see what it can do...
7
28
2,372
update
1
3
409
I feel sorry for UK. They published "rules" on how to use AI: Basically saying to use AI only when needed, with short prompts and as few interactions as possible to reduce environmental impact 🤣 Will probably be adopted by EU too.
If people think this is a joke - it’s not: ai.gov.uk/knowledge-hub/how-…
33
11
185
9,402
Useless in what? For being completely private? For having no limits? For not being nerfed after a while? For having no downtime? For being able to own intelligence instead of renting one? Yes, without all of these local models are completely useless I agree 💯
Local models are useless
41
9
181
7,210
After the new additions to the Qwen3.8-Flash-Next for a single spark, coming with similar updates for the dual recipe! Testing rn to ensure it's stable. Dropping possibly tomorrow!
Qwen3.8-Flash-Next on a single @NVIDIAAI DGX Spark just got better 🚀 Default recipe 👇 ・Fixed a hidden bug that dropped per-layer embeddings: quality loss 1.397 → 1.344, same speed ・Follow-up replies 2x faster (1.66s → 0.81s), after tool calls 3.61s → 2.72s ・~49 tok/s prose, ~62 code (1 stream), up to ~168 tok/s (4 streams) ・Optional bit-exact decoding for evals and debugging ・Speculative depth 6 unlocked, +16% on code New opt-in lane on vLLM 0.30 (./start-v030.sh) 🌟 ・NVIDIA's official weights now fit on one Spark: the 48 GB per-layer table is read by the GPU straight from a file, built once and reused on every reboot ・FP8 KV backport: 801k tokens of context, 3/3 needles at 200k ・Best quality yet: NLL 1.332 ・Prefill +10% (2,140 tok/s), follow-ups 0.67s, code +6% at 4 streams ・Trade-off: prose ~25% slower. Use it for agents, keep the default for raw speed
Made with AI
7
2
35
4,200
Great article by @exolabs, @0xSero, and @alexocheema Recommend read for anyone who plans getting into local AI, and owns or plans to own one or more DGX Sparks. DGX Sparks are great little devices that punch well above their weight.
7
15
197
17,552
5
1,468
First time running Opus 5.5 with Ultracode Let's see what it can do...
17
1
131
10,729
Mia retweeted
Deepseek V4.1 Flash TP4 on 4 DGX sparks is fire. thank you @majewskizby @MiaAI_lab
4
3
32
3,145
In few days OpenAI will adopt this 100%. Claude will now try to find a "stopping point" if you reach the 5h limit, instead of just dying in the middle of a running task. Only allowed once per week for Pro sub ($20), but unlimited times for Max subs.
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
14
2
161
17,429
Opus 5.5 is insane, and can do incredibly complicated things if you just let it.
Gave Opus 5.5 donald's prompt, Midjourney, and a moodboard 12 hours later, woke up to this:
5
2
87
10,820
Be nice or leave. Got tangled up watching the online debates over open source models. Not me or my work.. Yet.. But, Here's where I landed. I'm building, learning and growing. Please do the same. This is my "be nice or leave" sign. Copy it. Use it. It was mine, now it's yours. No one owns it, because all knowledge is by definition distilled. It doesn't belong to anyone. I hand-coded in a trailer before most of you were born. Those days are over. Information is moving too fast for anyone to claim they got there first. Move on. Less talk more building. If you wrote the code, show me in a notebook. Otherwise, be nice or leave. Copy the sign. Use it. No debate with the actors.. post the sign.. Handle changes are coming.. seen this episode before. Be nice or leave..
1
9
1,067
10.7M impressions in the last two weeks 😲 Thanks to everyone who read, liked, shared, and replied 🫶 If you want to know how I did it, it's all in my book. mia-ai.net/book
25
6
241
7,464
buymeacoffee.com/ashhart I supported Ash's MCDMA/Drift Repo... Watch this man fly.. Going places.. Mark my words... ! @ashxhart a must follow....! This man speaks RDMA as a second language @NVIDIAAI... @alexocheema please tell me your talking to this guy!! I cannot see a @exolabs @nvidiaai without him...
3
2
8
3,908
Mia retweeted
X is in league of its own. I’ve met some incredibly generous people investing in open source work, had three job offers and seen some absolutely phenomenal work being done by others. I genuinely look forward to seeing what is cooking. Thank you to everyone who is contributing to the open source community. @volatilemarkts @MiaAI_lab @Kurcide
5
1
28
1,817
Models now allows you see different performance stats depending on # of DGX Sparks. I've also added prefill numbers. mia-ai.net/models
19
5
161
10,833
What's your go-to model on your DGX Sparks? Also comment why. I'd like to understand which models are being used the most so I know where to focus my time on improving and optimizing performance.
38% GLM 5.3 Flash
18% DeepSeek v4.1 Flash
33% Qwen3.8-Flash-Next
11% Other (comment)
2,342 votes • 2 hours
131
2
100
19,678
Interesting to see DeepSeek v4.1 Flash only at 3rd place
6
6
1,404
Look at what is possible, and try to think on where things are going for local AI. We will do everything we can to make local AI shine.
Pushing this further, I loaded @NVIDIAAI's Nemotron Lightning into the TensorFold engine on my M5 Max MBP and achieved over 200tps decode on the first run. There is headroom here too, I could see 250/60 achievable. Qwen 3.8 Flash Next is up next :P
6
2
82
9,775
Coming soon: Qwen3.8 Flash Next update for a solo DGX Spark ⚡️
The latest improvements on vLLM enhances serving @Alibaba_Qwen Qwen3.8 Flash Next. Improvements are coming to the recipe!
19
15
330
21,671
Now LIVE
A good day for updates ⚡️ Update for Qwen3.8 Flash for a solo DGX Spark - Decode on prose is now 49 tok/s single stream. - Decode on code is now 62 tok/s single stream. - Faster follow-up replies. - Optional vLLM 0.30 path with official Nvidia NVFP4. This is still the BEST model to run a single DGX Spark.
6
1,213