Principal Engineer, Founder @AICTechPro | Obsessed with local AI & MLX on Apple Silicon.

Massachusetts, USA
Dual Mac Studio rack. One Thunderbolt 5 cable. Plug-and-play hardware with its own console. We’ve been building the software for the last few months: production local inference on open weights (distribution, cache, speculation, tool loops, API surface). Private AI hardware + full production software. If you’re interested, reach out. Video soon.
1
7
220,690
First @typesafeai use case, live in our Mac app: setup and troubleshooting help when no model is loaded. Model downloading, load failed, API returning 503, phone won't pair: the user asks, Jev reads the question with the whole built-in manual as state and decides, with probabilities, what it is and which article answers it, or that nothing does. The app then shows the real documentation and live status. Jev decides, the app answers from its own docs. No model loaded, nothing invented. 42/42 on a held-out set: paraphrases, typos, French, German, Spanish, features that don't exist, follow-ups. Median 0.93 s. Great breakthrough by the TypeSafe team. Thank you.
1
1
5
543
We got access to TypeSafe and Jev. Thank you to the @typesafeai team and @CompleteSkeptic for the access and for the breakthrough itself: a model that returns calibrated decisions instead of generated text is a new primitive to build with. Our first use case: an on-device request router. Today it asks the Apple Foundation Model nine yes/no questions plus a pick from a help catalog, in two passes. With Jev that becomes one request, every question in parallel, each answer a probability instead of a hard boolean. We're starting with it as the judge for our routing evals. Next: demos of every use case, routing, ranking, extraction, verification and guardrails, in our agentic app. Follow-ups coming.
4
293
DeepSeek-V4.1-Flash, fully local. No cloud. 2× Mac Studio M3 Ultra (512 GB each), one TB5 cable (RDMA ~8.7 GB/s), TP=2. ~267 GB resident per box. 748B total (552B backbone + 196B Engram). 8B active prefill / 16B decode. 1M context. Image in. 40–66 tok/s decode with DSpark (70–81% accept). ~1,100 tok/s prefill. Sub-second TTFT on cached agent turns. Our MLX port. Weights stay in the room.
6
11
568,178
Prompt cache on our headless dual Mac Studio cluster was stuck at 0%. Silent anomaly. Agent turns re-prefilled 12k–29k tokens every step (56% of a 10 min session). Two server fixes → 99/100/100% hits, 18.5 s → 12.0 s per turn. Fix #1: make the cache key stable across tool loops. Fix #2: a one-token warm so the next turn actually hits. If your local agent feels mysteriously slow, check cache hits before you blame the hardware.
2
3
143
15.6 tok/s is not the story. Stability and task quality are. 1M context on a production-shaped local server. Next: speculative decoding for more speed, and the cluster hardware video.
GLM-5.3 (753B), 8-bit mxfp8 · 2× M3 Ultra · TB5 RDMA: 411 GB/rank, 15.6 tok/s. Do not be fooled by tok/s. Most stable, most capable local model we tested. Agent tasks done well. 1M context. Server optimized for production inference. Cluster hardware video next. Stay tuned.
1
1
3
189
Ran this on the dual M3 Ultra rack. Speculative decoding isn’t magic if the drafter isn’t in the checkpoint. The 14.5 → 42 jump was.
14.5 → 42 tok/s on DeepSeek-V4-Flash. Same weights (87 GB/rank). Same 2× M3 Ultra over TB5 RDMA. Only change: the drafter that ships in the checkpoint (78% accept). The cycle cost of the other 22% is next. Follow if you want that breakdown.
2
138
There's growing talk, in more than one capital, of limiting who can download open source model weights. It's like telling early developers that programming languages were too dangerous to share. We shared them anyway, and it built the modern world. Labs made the breakthroughs. Home desks will carry them further. History repeats, and openness wins.
2
2
76
I asked DeepSeek-V4-Flash-Vision-Exp to look at my webcam. It recognized the laptop on my desk was running the exact same DeepSeek Harness UI it was using to reply to me. It saw its own reflection. It saw the rack with the Mac Studios too. From the Harness custom memory I set up, it knew that was the local MLX server it was running on. The loop is real. 🌀 I like where this is going.
2
3
185
Asked GLM 5.3 on MLX to organize my files by content. No image input, so read_image was off the table. So it wrapped macOS Vision in a shell call and piped the text back to itself. It built itself an eye.
1
2
7
346
Day 0 open weights, full BF16, zero quantization, on Apple Silicon via MLX: • ONE Mac Studio (M3 Ultra, 512 GB): 26.9 tok/s sustained decode • TWO Studios, tensor-parallel over Thunderbolt 5: 33.8 tok/s Their disclosed architecture is what makes it possible: • 6B active of 125B: per-token compute fits consumer silicon, even at 360 GB of BF16 • 51B N-gram table is a lookup, not a matmul. We keep it memory-mapped, a few KB pages in per token. 100 GB of RAM saved. • The in-checkpoint MTP head drafts for the base model: self-speculative decoding at ~70% acceptance, no separate draft model • GDN + QSA keep the cache tiny at long context Thank you @Alibaba_Qwen for the amazing work.
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: qwen.ai/blog?id=qwen3.8-flas… - Technical Report: github.com/QwenLM/Qwen3.8-Fl… - Hugging Face: huggingface.co/Qwen/Qwen3.8-… - ModelScope: modelscope.cn/models/Qwen/Qw…
1
3
219
What a moment for local AI. M5 Ultra is out. Same 512 GB as the dual M3 Ultra cluster we already run, now with 1.2 TB/s (+50%) and Neural Accelerators on every GPU core. Apple claims ~4× LM Studio prompt vs M3 Ultra. Amazing hardware.
1
3
203
Hit 42 tok/s locally on DeepSeek-V4-Flash with DSpark speculative decoding on MLX (2× Mac Studio M3 Ultra ). We've been deep in MLX + production server optimization and custom kernels for a while now. Anyone else pushing on MLX optimizations? What are you benchmarking? Curious what others are seeing 👀
1
5
241
This is a great example of Think Different. More breakthroughs are coming that will change AI compute.
GPUs don’t have to do all the work. Pairing GPUs with FPGAs for long-context inference can play to the strengths of each architecture and improve efficiency. Jason Cong, Distinguished Professor of Computer Science at @UCLA, explains
1
1
139
Intelligence is the model plus the harness. Just watched @deepseek_ai Harness catch a loop. Same failed tool call, three times. Then it injected: analyze the last result, change the arguments or change the approach. That’s the stack doing real work.
2
5
249
Don't fall into the trap of just waiting for smarter models. If you're a developer and obsessed with AI, there's a lot you can do. Inference frameworks, harnesses, and the rest of the AI software stack still need improvements and breakthroughs. Let's keep building.
2
3
84
Elon told the SpaceX team they have to win AI in hardware and in software. Hardware they already know how to build. Gigawatts, clusters, later the sky. Software was the missing piece. That is why Cursor is inside SpaceX now. The team that knows what you do with the model after it exists. Grok Bot is that layer shipping.
1
231
Malek Ould-Oulhadj retweeted
This is why we’re building Izuran. An app that lives on that desk.
Our current @bot workflow: the Mac Studio is the working machine for Grok Bot. Local models, Python jobs, the agent flows. It stays on in the office. I send work from the phone. 512 GB on the desk. Grok Bot runs the loop on that machine instead of shipping the jobs out.
2
2
38
Malek Ould-Oulhadj retweeted
Grok Bot on the Mac Studio is the first time this loop feels like work, not a demo 👏
Our current @bot workflow: the Mac Studio is the working machine for Grok Bot. Local models, Python jobs, the agent flows. It stays on in the office. I send work from the phone. 512 GB on the desk. Grok Bot runs the loop on that machine instead of shipping the jobs out.
2
3
93
Our current @bot workflow: the Mac Studio is the working machine for Grok Bot. Local models, Python jobs, the agent flows. It stays on in the office. I send work from the phone. 512 GB on the desk. Grok Bot runs the loop on that machine instead of shipping the jobs out.
1
3
293