Part 2 of my local AI test is live. Part 1 was a few days ago: 5 setups on 4 memory classes, 10 app prompts, 45 minutes each. It's still on the page. Part 2: 9 setups, same 4 classes, same 10 tasks. New: kits and recipes from @MiaAI_lab and @ashxhart, and the 3 biggest tasks get up to 150 minutes. 16 GB (RTX 5060 Ti) · Qwen3.8-27B Mia's exllamav3 kit and my llama.cpp 32 GB (RTX 5090) · Qwen3.8-27B Mia's vLLM recipe and Ash's TensorFold 128 GB (GB10) · Qwen3.8-Flash-Next Mia's TensorFold recipe and Mia's vLLM kit 256 GB (2× GB10) · GLM-5.3-Flash Mia's vLLM, Mia's TensorFold, Ash's TensorFold Every app they built is clickable, task by task. Details in the replies. nipale-ai.github.io/local-ai…
6
6
589
Replying to @MiaAI_lab @ashxhart
128 GB (one GB10 box each), model Qwen3.8-Flash-Next, both are Mia's recipes: TensorFold 0.3.6.3, MLX 4bit 105.5 tok/s, TTFT 4.99 s (first request of a task), 0.82 s later vLLM kit, NVFP4 72.8 tok/s, TTFT 2.12 s (first request of a task), 1.21 s later tok/s = output tokens per second per request, median. TTFT = p50. Apps: nipale-ai.github.io/local-ai…

Oct 3, 2026 · 6:04 PM UTC

1
35
Sort replies: Relevant Recent Liked