CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering tinyurl.com/7ymyv4sb | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡

PNW
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. ๐Ÿ’ก I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe ๐Ÿ‘‰ ๐Ÿง  Qwen3.5-4B Q4_K_M GGUF โš™๏ธ Ryzen 5 7540U โ€” 6C/12T ๐Ÿงต --threads 9 ๐Ÿงต --threads-batch 12 โšก --prio 2 ๐Ÿ”„ --poll 50 ๐Ÿ“ฆ --batch-size 2048 ๐Ÿ“ฆ --ubatch-size 512 ๐Ÿš€ --flash-attn on ๐Ÿง  KV cache: q4_0 / q4_0 ๐Ÿ”ง --repack ๐Ÿ’พ --mmap ๐Ÿ‘ค --parallel 1 ๐Ÿšซ --device none ๐Ÿšซ --gpu-layers 0 ๐Ÿšซ KV/op GPU offload ๐Ÿšซ MTP OFF Interesting result ๐Ÿ‘‰ MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp โ€” on a basic laptop CPU.
11
1
47
29,010
You can put Codex + Claude Code + Qwen Code + OpenHands behind one API. That's GitHub project HarnessRouter. Instead of building separate integrations for every coding harness, your app talks to HarnessRouter and it handles the different backends. ๐Ÿค– Codex ๐Ÿง  Claude Code ๐Ÿˆ Qwen Code ๐Ÿ’ป OpenCode ๐Ÿ› ๏ธ Hermes ๐Ÿฆพ OpenHands & more ๐Ÿ‘‰ behind one API. But they fixed a pretty funny problem. ๐Ÿ˜‚ In one reproduced Codex failure: ๐ŸŒ Before 28 calls 27 sleeps 321.9 seconds โ†’ generic error ๐Ÿš€ After 1 call 0 sleeps โ†’ actual error immediately What happened? Codex rejected the request with a 400. But HarnessRouter could mistake that rejection for something temporary...so it tried again...and again. 28 times. For more than 5 minutes. ๐Ÿ˜‚ v0.23.12 fixes it. A deterministic non-429 4xx rejection is now basically ๐Ÿ‘‰ Nope. The runner rejected it. Stop. Things that actually might recover 429s, startup responses and 5xx errors can still retry. Just one AI harness learning that when Codex says no...maybe don't ask it another 27 times. ๐Ÿ”— Link in ALT
14
2
10
1,196
These two GPUs are almost a decade old and not even the same size but they're getting ~40 tps on Qwen3.8-27B - beating stock Q6 on Strix Halo ๐Ÿ˜‚ Someone grabbed ๐ŸŽฎ Tesla V100 16GB ๐ŸŽฎ Tesla V100 32GB 48GB of mismatched old Nvidia VRAM... and got Qwen3.8-27B Q6 running at: ๐Ÿš€ 39.88 tps generation โšก 1,377 tps prompt processing Even with a 16K prompt it was still processing at ~1,221 tps. ๐Ÿ‘€ No exotic inference engine either. It's ordinary mainline llama.cpp + tensor splitting. They loaded the much bigger Qwen3.8-Flash-Next MoE onto the same two old V100s. That ran at ๐Ÿง  26.29 tps On two mismatched GPUs released in 2017. This is why I keep watching old datacenter hardware going for crazy prices on Ebay. Everyone is chasing the newest GPUs... Meanwhile people keep finding ways to make yesterday's discarded server hardware run Local AI. Now I want to know what 48GB of used V100 VRAM actually costs. ๐Ÿ˜‚ ๐Ÿ”— Link in ALT
4
1
15
1,169
OpenRouter recently published this chart. No it isn't your imagination, model releases are getting faster and nearly impossible to stay on top of.
2
3
10
727
This guy was getting ~15 tps running Qwen3.8-Flash-Next on his 12GB RTX 5070. Apparently that wasn't good enough. ๐Ÿ˜‚ So he built his own inference engine. (as we all should) Now he's reporting ๐ŸŒ llama.cpp โ†’ ~15 tps ๐Ÿš€ Strata โ†’ up to 65.1 tps And this isn't big workstation either Specs ๐ŸŽฎ RTX 5070 12GB ๐Ÿง  64GB DDR5-5600 โš™๏ธ Ryzen 5 7600 ๐ŸชŸ Windows At 128K context his new Strata engine reports: Q2_0 โ†’ 65.1 tps + 543 tps prompt IQ2_XS โ†’ 52.0 tps + 472 tps prompt IQ3_XXS โ†’ 44.8 tps + 414 tps prompt For Qwen3.8-FLASH-NEXT. ๐Ÿ‘€ He built it specifically around this model and this kind of CUDA + system-RAM setup, paired with RCO-GSQ quants. ๐Ÿ‘‰ And he open-sourced it. I love these kind of Local AI projects. ๐Ÿ”— Link in ALT
53
90
1,185
66,558
This really matters for Local AI! We keep shrinking the model weights. Now researchers are going after the model's working memory. A new technique called MILO cuts KV-cache memory by up to ๐Ÿ‘‰ 50% ๐Ÿ‘ˆ and reports up to 1.8ร— higher throughput on Qwen2.5 experiments. ๐ŸŽฏ This is really signficant because KV cache grows with context. So once the quantized model fits into VRAM/RAM, context is the next memory hog. As an example ... ๐Ÿง  8GB KV cache @ 128K context โ†“ 50% compression ๐Ÿง  ~4GB for the same cached context So that saved memory could mean ... ๐Ÿ“š ~2ร— the context for the same KV-memory budget ๐Ÿค– more simultaneous agents/sessions ๐Ÿ’พ less KV spilling into slower RAM ๐Ÿง  more room for model weights MILO breaks the KV cache into blocks, finds redundancy using low-rank compression, and gives information-dense blocks more precision while compressing redundant blocks harder. โš ๏ธ MILO currently targets many-shot long-context inference and was tested on Qwen2.5 3B/7B with an Nvidia L4. Obviously this is research and not a llama.cpp feature you can turn on tonight. ๐Ÿ”— Paper in ALT
12
17
77
2,747
Thunderobot is turning Strix Halo mini PCs into a home AI rack. ๐Ÿ‘€ Their new 2-liter Station has 128GB of unified memory. But here's the part that caught my eye, it has ... ๐ŸŒ 2x 10GbE ports. They're pitching these small machines for multi-node AI clusters. So instead of ๐Ÿง  1 mini PC โ†’ 128GB They're now like ... ๐Ÿง  2 nodes โ†’ 256GB ๐Ÿง  4 nodes โ†’ 512GB And this isn't theoretical. We've already seen 4ร— Ryzen AI Max+ 395 mini PCs combine 512GB of memory to run DeepSeek V4 Flash locally. ...and now hardware makers appear to be designing around that idea. The "Station" itself is ๐Ÿ“ฆ 2 liters ๐Ÿง  Ryzen AI Max+ 395 ๐Ÿ’พ 128GB LPDDR5X ๐ŸŽฎ up to 96GB for graphics โšก 132W ๐ŸŒ dual 10GbE ๐Ÿ”Œ dual USB4 ๐Ÿ’ฝ 3ร— M.2 Thunderobot says a Ryzen AI Max+ PRO 495 version is coming. That platform supports: ๐Ÿง  192GB unified memory ๐ŸŽฎ up to 160GB for graphics Damn, 4x 192GB nodes? That's 768GB sitting on your desk at home. ๐Ÿ‘€ โš ๏ธ Ethernet doesn't turn four PCs into one 768GB computer. Distributed inference needs software designed to split the workload across nodes like the approaches we're starting to see with DwarfStar, llama.cpp RPC and Cascadia. And the current 128GB Station reportedly costs around $3,800โ€“$4,000 in China. ๐Ÿ”— Link in ALT
11
2
28
5,802
Okay this is a weird big LLM launch. Why so low-key?๐Ÿ˜‚ Meituan dropped LongCat-2.5-Preview. Stats ... ๐Ÿง  1.6 TRILLION parameters โšก ~48B active/token ๐Ÿ“š 1 MILLION token context ๐Ÿ‘€ Native multimodal ๐Ÿค– Built for long-running agents That's all we know. You can already plug it into Claude Code, OpenClaw, OpenCode, Hermes and Kilo Code. The price is kinda nuts too: ๐Ÿ’ฐ $0.30 / 1M input tokens ๐Ÿ’ฐ $0.006 / 1M cached input ๐Ÿ’ฐ $1.20 / 1M output But here's the strange part... where are the benchmarks? I went looking. Nope. No benchmark table. No LongCat-2.5 Hugging Face weights I can find. No GGUF. No model card with the usual wall of scores. That's especially interesting because LongCat-2.0 was open-weight. So for the moment, Meituan is handing us a 1.6T multimodal agent with a 1M context window and saying ๐Ÿ‘‰ here, try it. Why not codename on OpenRouter? And apparently existing users just got 5 million free tokens to do exactly that. ๐Ÿ”— Link to X post in ALT
18
4
101
7,099
Someone found ~900 more prompt tps hiding inside the RX 7900 XTX. ๐Ÿ‘€ Same RX 7900 XTX + Gemma 4 26B-A4B Q4_0 now.. ๐ŸŒ 3,410 tps โ†’ ๐Ÿš€ 4,331 tps That's ~27% faster prompt processing. Token Gen got a smaller bump too 135.6 โ†’ 142.8 tps So what changed? llama.cpp. They just merged a new INT8 cooperative-matrix Vulkan path specifically for AMD RDNA3 + RDNA4 GPUs. and it covers a bunch of the quants we actually run ... ๐Ÿง  Q3_K / Q4_K / Q5_K / Q6_K ๐Ÿง  IQ4_XS ๐Ÿง  MXFP4 ๐Ÿง  NVFP4 ๐Ÿ”— Link in ALT
3
2
44
3,136
llama.cpp may be about to significantly improve PP on Intel Arc. New Vulkan Flash Attention kernel. On an Intel Arc Pro B70 32GB / Windows: ๐Ÿง  Qwen3 8B Q4 1,065 โ†’ 2,956 prompt tps ๐Ÿš€ โ‰ˆ 2.8ร— faster ๐Ÿง  Qwen3.8-27B Q4 659 โ†’ 932 tps โ‰ˆ 41% faster And on Arc B390: ๐Ÿง  Qwen3 8B Q4 235 โ†’ 912 tps ๐Ÿ‘‰ almost 3.9ร— faster Nothing about the GPU changed. The software learned how to use it better. This is why Local AI hardware comparisons can become outdated quickly. A GPU that looks mediocre today can suddenly become better when llama.cpp gets a better kernel. โš ๏ธ These are PP gains, not 3ร— faster token generation. โš ๏ธ PR is currently open / not merged. ๐Ÿ”— Link in ALT
1
5
49
3,829
Wow, speculative decoding just gave a local vision model up to a 3.13ร— decode speedup. Liquid AI released LFM2.5-VL-3B-DSpark. The idea is speculative decoding for vision ๐Ÿง  Main model: LFM2.5-VL-3B โšก Draft model: ~280M parameters ๐Ÿ“ˆ only +8.9% more parameters And Liquid reports ... ๐ŸŽ M5 Max / MLX-VLM ๐Ÿš€ 2.30โ€“3.13ร— faster decoding โšก up to 2.62ร— better end-to-end ๐ŸŽ M3 Ultra / llama.cpp ๐Ÿš€ 1.57โ€“2.14ร— faster decoding โšก up to 1.77ร— better end-to-end DSpark is basically ... The model guesses several tokens ahead. The full 3B model verifies them together. If the guesses are right โ†’ you avoid repeatedly hauling the whole big model through memory for every single token. And because the full model verifies the draft, the output distribution stays the same as the target model under matched sampling. Even better for Local AI: โœ… llama.cpp โœ… GGUF โœ… MLX-VLM โœ… SGLang โš ๏ธ Liquidโ€™s published tests used FP16/BF16, so donโ€™t assume the same 3.13ร— gain for Q4 GGUFs yet. ๐Ÿ”— Link in ALT
7
4
25
2,570
Is this the Qwen3.8-27B fine-tune everyone has been asking for? It has the two fixes people keep asking for in Qwen3.8-27B. (catchy name: ThinkingCap-Qwen3.8-27B-Uncensored-Heretic-GGUF) Less overthinking + fewer refusals. Stats ... ๐Ÿง  Qwen3.8-27B โ†’ ๐ŸŽ“ ThinkingCap โ†’ ๐Ÿ”“ Heretic โ†’ ๐Ÿ“ฆ GGUF ThinkingCap reports... โšก 37.2% fewer thinking tokens ๐Ÿ“Š 86.6 โ†’ 85.8 avg accuracy And an independent llama.cpp Aider test today found: Vanilla Qwen: ๐Ÿง  12,547 median tokens โฑ๏ธ 1,481 sec/case ThinkingCap: ๐Ÿง  7,436 tokens โฑ๏ธ 777 sec/case โ€ฆwith the SAME 77.6% retry-pass score in that test. Then Heretic removes most refusal behavior. Now mradermacher has local GGUFs: ๐Ÿ“ฆ IQ4_XS โ€” 15.3GB ๐Ÿ”ฅ Q4_K_M โ€” 16.6GB ๐Ÿ’Ž Q6_K โ€” 22.2GB Runs on ... โœ… llama.cpp โœ… LM Studio โœ… Ollama โœ… multimodal โœ… local agents So this is basically ๐Ÿ‘‰ Qwen3.8-27B that wastes fewer tokens, refuses less, and fits on normal Local AI hardware. ๐Ÿ‘€ โš ๏ธ Important: ThinkingCap uses a PolyForm Small Business license, not Qwen's original Apache-2.0 license. ๐Ÿ”— Link in ALT
7
7
104
6,109
Those four LoRAs aren't necessarily the end of the model for the model. Mind Lab designed the architecture so you can add on ๐Ÿง  keep ONE frozen base model โž• train a new specialist LoRA ๐Ÿ”Œ register it with the router ๐Ÿ”„ upgrade it independently โช roll it back independently ๐Ÿšซ remove it without touching the base They even say teams/users could mount their own specialist adapters alongside the built-in ones.
This new 752B agent model has an interesting Local AI architecture. Macaron-V1.1 isn't four enormous specialist models but it has specific LoRAs. ๐Ÿง  744B GLM-5.3 base ๐Ÿ’ฌ + 2B Chat LoRA ๐Ÿค– + 2B Agent LoRA ๐Ÿ’ป + 2B Coding LoRA ๐ŸŽจ + 2B GenUI LoRA One giant shared base intelligence w/ four relatively tiny specialist adapters. And GLM-5.3 itself is a MoE with only ~40B parameters active/token. Instead of storing separate Coding + Agent + Chat models keep 1 base model locally and swap small specialist brains onto it. Macaron is doing it with 1.1 at ridiculous 752B scale. Their smaller Macaron-Tall already applies the idea to a Qwen 35B-A3B base intended for local deployment. ๐Ÿ‘€ โš ๏ธ V1.1 itself is still enormous, and I haven't found an official V1.1 consumer quant/local benchmark yet.
2
4
1,232
I've been posting about MediaTek and how they claim a phone can run a 30B AI model? Qualcomm went one step further. ๐Ÿ‘‰ They actually demo'd one. At Snapdragon Summit, Qualcomm + partners ran StepEdge-Omni 30B-MoE directly on a Snapdragon 8 Elite Extreme Gen 6 reference design. ๐Ÿง  30B total parameters โšก ~3B routed params active/token ๐Ÿš€ >330 tok/s prefill ๐Ÿ—ฃ๏ธ >28 tok/s decode ๐Ÿ’พ >50% lower runtime-memory needed ๐Ÿ’ฝ model weights can load from phone flash โ˜๏ธ no cloud inference required for the demonstrated agent workflow And this wasn't just โ€œask the LLM a question.โ€ The local agent demonstrated: ๐Ÿ“ง understanding email โœˆ๏ธ planning trips ๐Ÿ“… updating calendars ๐Ÿจ flight/hotel recommendations ๐Ÿ“จ drafting replies So they are formalizing the memory hierarchy too โšก CPU/GPU/NPU โ†’ compute ๐Ÿง  RAM โ†’ active/hot model data ๐Ÿ’ฝ phone flash โ†’ model weights We keep seeing the same Local AI architecture move downward the hardware hierachy ... GPU โ†’ RAM โ†’ SSD and nowโ€ฆmobile compute โ†’ phone RAM โ†’ UFS flash. ๐Ÿ‘€ โš ๏ธ This was a Qualcomm reference-platform demo, not yet a benchmark from a retail phone. We have to put Qwen on the production hardware and give me tps. ๐Ÿ˜‚ ๐Ÿ”— Link in ALT.
5
12
1,090
Hugging Face did Apple silicon a solid and users can use GGUFs directly. ๐ŸŽ ๐Ÿคฏ Transformers can now load GGUF quants directly on Apple Silicon and keep the weights packed instead of expanding them back to BF16. It even reuses ggml Metal kernels underneath. On an M2 Max ๐Ÿง  Qwen3.5-4B Q4 Transformers โ€” 70.4 tok/s llama.cpp โ€” 71.8 ๐Ÿง  Qwen3.8-27B Q4 Transformers โ€” 15.9 tok/s llama.cpp โ€” 13.4 ๐Ÿง  Qwen3.5-35B-A3B IQ4 Transformers โ€” 60.2 tok/s llama.cpp โ€” 61.3 So now the same compact GGUF can live inside the normal: ๐Ÿ Python ๐Ÿ”ฅ PyTorch ๐Ÿค— Transformers ecosystem without giving up llama.cpp-style quantization. Does this kill MLX? No. MLX is still purpose-built for Apple Silicon. But it means Mac Local AI developers have another near-llama.cpp-speed path without leaving Transformers. ๐Ÿ‘€ โš ๏ธ Initial packed-GGUF support is Apple Silicon + Qwen3.5/3.8 focused, and HF says the benchmark conditions arenโ€™t perfectly identical. ๐Ÿ”— Link in ALT
6
20
195
13,011
๐Ÿ˜‚ Well that didn't take long. Yesterday when I posted Apple's new LensVLM-9B I had ... โ“ llama.cpp โ€” not yet / not verified 24 hours later boomโ€ฆ โœ… llama.cpp โœ… GGUF โœ… LM Studio โœ… Ollama โœ… MLX Bartowski already has the 18.8GB BF16 model squeezed down to: ๐Ÿ”ฅ Q4_K_M โ€” 5.84GB ๐Ÿค IQ2_M โ€” 3.54GB the multimodal projector for image/document input. So Apple's weird Qwen-derived document AI went to a ~6GB GGUF and can now be used on normal Local AI hardware. ๐Ÿ‘‰ in basically a day. ๐Ÿ˜‚ This is why open/downloadable weights matter. The community immediately takes over. ๐Ÿ‘€ ๐Ÿ”— Link in ALT.
Hey Apple, why did you copy Qwen? ๐Ÿ™‚ I have to download this (will it run on CUDA?) Apple quietly dropped a 9B downloadable AI model on Hugging Face. And it is Local AI friendly handling huge documents. LensVLM-9B is Appleโ€™s post-trained Qwen3.5-9B model. Instead of stuffing a giant document into the context window, this model โ€ฆ ๐Ÿ“„ render pages as compressed images ๐Ÿ‘€ scan them cheaply ๐Ÿ”Ž find the pages that matter ๐Ÿ”ฌ expand ONLY those pages ๐Ÿง  answer from the full-resolution content Apple says it maintains near full-text accuracy at 4.3ร— effective compression and beats other compression/retrieval approaches up to 10.1ร—. And itโ€™s small enough to run locally ๐Ÿง  9B params ๐Ÿ’พ 18.8GB BF16 โœ… downloadable weights โœ… Transformers โœ… vLLM โœ… SGLang โ“ llama.cpp โ€” not yet / not verified This is an interesting way to attack a Local AI bottleneck. โš ๏ธ Built on Qwen3.5-9B-Base and Apple post-trained VLM, not a new Apple foundation model. ๐Ÿ”— Link in ALT.
2
3
28
4,747
Qwen's brand-new image model is getting GGUFs (and uncensored) ๐Ÿซฃ Qwen-Image-2.1 only dropped a few days ago and its 7B image generator normally weighs ๐Ÿ‘‡ ๐Ÿ˜ BF16 โ€” 14.23GB Now only ... Q8 โ†’ 7.59GB Q6 โ†’ 5.88GB Q5โ†’ 5.22GB Q4_K_M โ†’ 4.60GB ๐Ÿ‘ˆ๐Ÿ‘€ ๐ŸŽฎ keep the ~4.6GB image model in VRAM ๐Ÿง  put the ~9.35GB text encoder in SYSTEM RAM The text encoder only runs once per prompt, while the image model stays on the GPU for the expensive sampling. So you don't necessarily need a giant 24GB GPU to experiment with Qwen's newest image model locally (uncensored). And Qwen-Image-2.1 isn't just text-to-image but also ... ๐ŸŽจ generation + editing ๐ŸชŸ transparent RGBA images ๐Ÿ–ผ๏ธ up to 10 reference images โœ๏ธ local/masked editing Runtime support is already broad ... โœ… ComfyUI โœ… ComfyUI-GGUF โœ… stable-diffusion.cpp โœ… Diffusers โœ… SGLang โœ… vLLM-Omni โœ… CUDA / Vulkan / Metal One strange wrinkle โ†’ The uncensored GGUF isn't actually an abliterated fine-tune. It's the original upstream Qwen weights. Can anyone explain?? ๐Ÿ”— Link in ALT
6
10
178
13,500
๐ŸŽ‰ This is a pretty big win for AMD Local AI. ๐ŸŽฏ Perplexity Portable Computer is coming to Ryzen AI Halo. And this isn't another DIY pile of scripts. It's a prepackaged local agent stack being brought to AMD's giant-memory AI PC platform. Portable Computer puts the stack on your PC ๐Ÿ‘‡ ๐Ÿง  local model ๐Ÿค– agent + orchestrator ๐Ÿ”Ž local search ๐Ÿ“ works with your local files/apps ๐Ÿ› ๏ธ tools + actions โฐ recurring workflows ๐Ÿ”’ sandboxed execution โ˜๏ธ cloud escalation only when you approve it Until now, Perplexity's documented local-inference hardware has been ๐ŸŸข DGX Spark ๐ŸŸข NVIDIA RTX PCs with 24GB+ VRAM Now AMD is bringing it to Ryzen AI Halo. So now ... ๐Ÿ’พ current Halo โ†’ 128GB unified memory ๐Ÿš€ next Halo โ†’ 192GB unified / up to 160GB VRAM This is what I've been waiting for from these giant-memory AMD machines. Not everyone wants to constantly tinker just to keep a local-agent stack running. The interesting part is making Local AI easy enough that you spend your time ๐Ÿ‘‰using ๐Ÿ‘ˆ the agent instead of maintaining it. Your models. Your files. Your workflows. On your machine. ๐Ÿ‘€ ๐Ÿ”— Link in ALT
2
2
7
1,678
GGUF quantization may be getting a size slider of sorts with a new FIT-GGUF quants! Normally, when downloading a local model, you choose a quant level Q3 โ†’ Q4 โ†’ Q5 โ†’ Q6 โ€ฆand hope it fits your RAM/VRAM budget. FIT-GGUF flips that around. Give it a target size or fidelity tier, and it can mix precision tensor-by-tensor to find a GGUF that meets that target. OpenBMB highlighted it on MiniCPM5-2B: ๐Ÿ“ฆ Quality โ€” 1.46 GiB โš–๏ธ Balanced โ€” 1.28 GiB ๐Ÿ—œ๏ธ Compact โ€” 1.21 GiB ๐Ÿค Mini โ€” 1.14 GiB Those labels represent progressively looser fidelity to the BF16 model, measured primarily with KL divergence and not different skills like coding vs. reasoning. And they run in normal llama.cpp / LM Studio. โš ๏ธ FIT does not claim its tensor allocation is universally optimal yet. ๐Ÿ”— Link in ALT
5
8
48
3,887
๐Ÿ”ŽBuried in NVIDIAโ€™s latest developer SDK updates are some big clues about what it expects RTX Spark Windows PCs to become. Nvidia added RTX Spark support to its In-Game Inferencing SDK, which will support ๐Ÿง  local AI inference ๐Ÿ‘ˆ๐Ÿ‘€ ๐Ÿฆ™ latest llama.cpp improvements๐Ÿ‘ˆ๐Ÿ‘€ ๐Ÿ’Ž Gemma 4 integration๐Ÿ‘ˆ๐Ÿ‘€ ๐ŸŽจ Stable Diffusion plugin๐Ÿ‘ˆ๐Ÿ‘€ ๐ŸŽ™๏ธ local speech AI๐Ÿ‘ˆ๐Ÿ‘€ ๐ŸชŸ ARM64 + x64 RTX Spark packages๐Ÿ‘ˆ๐Ÿ‘€ At the same time, RTX Kit is adding these game features ๐Ÿง  neural shaders ๐Ÿ—œ๏ธ neural texture compression ๐Ÿ”๏ธ Mega Geometry 2.0 ๐ŸชŸ more Windows ARM64 support Nvidia says NVIGI lets applications run AI models Locally, in-process, alongside the graphics workload. So imagine one PC doing: ๐ŸŽฎ graphics ๐Ÿ‘‰ ๐Ÿง  local LLM ๐ŸŽ™๏ธ speech ๐ŸŽจ image generation โ†’ simultaneously. This makes the 128GB RTX Spark Windows ARM machines so compelling. Nvidia looks like it is building an entire local AI + neural graphics PC platform around them. ๐Ÿ”— Link in ALT.
6
1
11
1,850
8 days after I posted about MediaTek claiming 30B Local AI on a phone, the first (maybe second) real phone using the chip is here (sadly, it's not available in the US yet). OPPO launched the Find X10 Pro Max with the 2nm Dimensity 9600 Pro. Inside this thing ๐Ÿ‘‡ ๐Ÿง  NPU 1090 โšก +51% LLM prefill ๐Ÿ”‹ +55% tokens/watt ๐Ÿงฉ on-device MoE support ๐Ÿ—œ๏ธ hardware KV-cache compression ๐ŸŽฎ new Arm Mali G2-Ultra NX GPU MediaTek says the NPU can support models up to 30B parameters. ๐Ÿ“ฑ This actual phone tops out at 16GB RAM ๐Ÿ’พ LPDDR5X, not LPDDR6 ๐Ÿ’ฝ UFS 4.1, not UFS 5.0 So we have the silicon. You know what we don't have yet!? โ“ 30B model running on it โ“ actual tok/s โ“ memory usage โ“ easy Hugging Face model loading Someone, please put Qwen on this thing. ๐Ÿ˜‚ Now I want the benchmarks!! ๐Ÿ™‚
1
12
1,526