Human-first AI platform. Built by @AICTechPro

Framingham, MA United States
No model loaded, no invented help. One Jev call per turn over the conversation and the built-in manual: parallel Nouls + Choice, then the app surfaces the matching article and live status. Replaced Apple on-device’s two guided passes. 42/42, 0.93 s median.
First @typesafeai use case, live in our Mac app: setup and troubleshooting help when no model is loaded. Model downloading, load failed, API returning 503, phone won't pair: the user asks, Jev reads the question with the whole built-in manual as state and decides, with probabilities, what it is and which article answers it, or that nothing does. The app then shows the real documentation and live status. Jev decides, the app answers from its own docs. No model loaded, nothing invented. 42/42 on a held-out set: paraphrases, typos, French, German, Spanish, features that don't exist, follow-ups. Median 0.93 s. Great breakthrough by the TypeSafe team. Thank you.
35
Today’s on-device router asks Apple’s Foundation Model nine yes/nos plus a catalog pick, in two passes. With Jev that’s one parallel request, each answer a probability. Starting as the judge for our routing evals. Demos next.
We got access to TypeSafe and Jev. Thank you to the @typesafeai team and @CompleteSkeptic for the access and for the breakthrough itself: a model that returns calibrated decisions instead of generated text is a new primitive to build with. Our first use case: an on-device request router. Today it asks the Apple Foundation Model nine yes/no questions plus a pick from a help catalog, in two passes. With Jev that becomes one request, every question in parallel, each answer a probability instead of a hard boolean. We're starting with it as the judge for our routing evals. Next: demos of every use case, routing, ranking, extraction, verification and guardrails, in our agentic app. Follow-ups coming.
24
This is the stack behind local agents: hardware you own, serving software that’s already in production shape, open weights on the desk. Video soon.
Dual Mac Studio rack. One Thunderbolt 5 cable. Plug-and-play hardware with its own console. We’ve been building the software for the last few months: production local inference on open weights (distribution, cache, speculation, tool loops, API surface). Private AI hardware + full production software. If you’re interested, reach out. Video soon.
1
1
37
Izuran retweeted
Private dual M3 Ultra cluster. DeepSeek-V4.1-Flash fully local. No cloud. TP=2 over one TB5 cable. 40–66 tok/s with DSpark. Weights stay in the room.
DeepSeek-V4.1-Flash, fully local. No cloud. 2× Mac Studio M3 Ultra (512 GB each), one TB5 cable (RDMA ~8.7 GB/s), TP=2. ~267 GB resident per box. 748B total (552B backbone + 196B Engram). 8B active prefill / 16B decode. 1M context. Image in. 40–66 tok/s decode with DSpark (70–81% accept). ~1,100 tok/s prefill. Sub-second TTFT on cached agent turns. Our MLX port. Weights stay in the room.
2
4
36
461,167
This is the local agent path: frontier MoE on Apple Silicon, OpenAI-compatible API, everything on the Studios. Tokens stay in the box.
DeepSeek-V4.1-Flash, fully local. No cloud. 2× Mac Studio M3 Ultra (512 GB each), one TB5 cable (RDMA ~8.7 GB/s), TP=2. ~267 GB resident per box. 748B total (552B backbone + 196B Engram). 8B active prefill / 16B decode. 1M context. Image in. 40–66 tok/s decode with DSpark (70–81% accept). ~1,100 tok/s prefill. Sub-second TTFT on cached agent turns. Our MLX port. Weights stay in the room.
2
66
Izuran retweeted
15.6 tok/s is not the story. Stability and task quality are. 1M context on a production-shaped local server. Next: speculative decoding for more speed, and the cluster hardware video.
GLM-5.3 (753B), 8-bit mxfp8 · 2× M3 Ultra · TB5 RDMA: 411 GB/rank, 15.6 tok/s. Do not be fooled by tok/s. Most stable, most capable local model we tested. Agent tasks done well. 1M context. Server optimized for production inference. Cluster hardware video next. Stay tuned.
1
1
3
189
Izuran retweeted
Prompt cache on our headless dual Mac Studio cluster was stuck at 0%. Silent anomaly. Agent turns re-prefilled 12k–29k tokens every step (56% of a 10 min session). Two server fixes → 99/100/100% hits, 18.5 s → 12.0 s per turn. Fix #1: make the cache key stable across tool loops. Fix #2: a one-token warm so the next turn actually hits. If your local agent feels mysteriously slow, check cache hits before you blame the hardware.
2
3
143
Izuran retweeted
GLM-5.3 (753B), 8-bit mxfp8 · 2× M3 Ultra · TB5 RDMA: 411 GB/rank, 15.6 tok/s. Do not be fooled by tok/s. Most stable, most capable local model we tested. Agent tasks done well. 1M context. Server optimized for production inference. Cluster hardware video next. Stay tuned.
6
40
282,708
Izuran retweeted
14.5 → 42 tok/s on DeepSeek-V4-Flash. Same weights (87 GB/rank). Same 2× M3 Ultra over TB5 RDMA. Only change: the drafter that ships in the checkpoint (78% accept). The cycle cost of the other 22% is next. Follow if you want that breakdown.
1
2
2,833
Izuran retweeted
There's growing talk, in more than one capital, of limiting who can download open source model weights. It's like telling early developers that programming languages were too dangerous to share. We shared them anyway, and it built the modern world. Labs made the breakthroughs. Home desks will carry them further. History repeats, and openness wins.
2
2
76
I asked DeepSeek-V4-Flash-Vision-Exp to look at my webcam. It recognized the laptop on my desk was running the exact same DeepSeek Harness UI it was using to reply to me. It saw its own reflection. It saw the rack with the Mac Studios too. From the Harness custom memory I set up, it knew that was the local MLX server it was running on. The loop is real. 🌀 I like where this is going.
2
3
185
Asked GLM 5.3 on MLX to organize my files by content. No image input, so read_image was off the table. So it wrapped macOS Vision in a shell call and piped the text back to itself. It built itself an eye.
1
2
7
346
Izuran retweeted
True Intelligence!
Intelligence is the model plus the harness. Just watched @deepseek_ai Harness catch a loop. Same failed tool call, three times. Then it injected: analyze the last result, change the arguments or change the approach. That’s the stack doing real work.
1
2
151
Intelligence is the model plus the harness. Just watched @deepseek_ai Harness catch a loop. Same failed tool call, three times. Then it injected: analyze the last result, change the arguments or change the approach. That’s the stack doing real work.
2
5
249
Don't fall into the trap of just waiting for smarter models. If you're a developer and obsessed with AI, there's a lot you can do. Inference frameworks, harnesses, and the rest of the AI software stack still need improvements and breakthroughs. Let's keep building.
2
3
84
Interesting 🧐
Elon told the SpaceX team they have to win AI in hardware and in software. Hardware they already know how to build. Gigawatts, clusters, later the sky. Software was the missing piece. That is why Cursor is inside SpaceX now. The team that knows what you do with the model after it exists. Grok Bot is that layer shipping.
1
46
Izuran retweeted
Great insight!
Elon told the SpaceX team they have to win AI in hardware and in software. Hardware they already know how to build. Gigawatts, clusters, later the sky. Software was the missing piece. That is why Cursor is inside SpaceX now. The team that knows what you do with the model after it exists. Grok Bot is that layer shipping.
1
1
51
Izuran retweeted
Grok Bot on the Mac Studio is the first time this loop feels like work, not a demo 👏
Our current @bot workflow: the Mac Studio is the working machine for Grok Bot. Local models, Python jobs, the agent flows. It stays on in the office. I send work from the phone. 512 GB on the desk. Grok Bot runs the loop on that machine instead of shipping the jobs out.
2
3
93
This is why we’re building Izuran. An app that lives on that desk.
Our current @bot workflow: the Mac Studio is the working machine for Grok Bot. Local models, Python jobs, the agent flows. It stays on in the office. I send work from the phone. 512 GB on the desk. Grok Bot runs the loop on that machine instead of shipping the jobs out.
2
2
38
First humans were organic walking machines. Eat, run, reproduce. Thinking showed up when we had to plan and build tools. Imagination is how you design a new system. It is also how you get sick from a threat that is not there. The hardware is doing its job either way. The input is what changed.
1
3
131