Part 2 of my local AI test is live.
Part 1 was a few days ago: 5 setups on 4 memory classes, 10 app prompts, 45 minutes each. It's still on the page.
Part 2: 9 setups, same 4 classes, same 10 tasks. New: kits and recipes from @MiaAI_lab and @ashxhart, and the 3 biggest tasks get up to 150 minutes.
16 GB (RTX 5060 Ti) · Qwen3.8-27B
Mia's exllamav3 kit and my llama.cpp
32 GB (RTX 5090) · Qwen3.8-27B
Mia's vLLM recipe and Ash's TensorFold
128 GB (GB10) · Qwen3.8-Flash-Next
Mia's TensorFold recipe and Mia's vLLM kit
256 GB (2× GB10) · GLM-5.3-Flash
Mia's vLLM, Mia's TensorFold, Ash's TensorFold
Every app they built is clickable, task by task. Details in the replies.
nipale-ai.github.io/local-ai…
6
6
589
128 GB (one GB10 box each), model Qwen3.8-Flash-Next, both are Mia's recipes:
TensorFold 0.3.6.3, MLX 4bit
105.5 tok/s, TTFT 4.99 s (first request of a task), 0.82 s later
vLLM kit, NVFP4
72.8 tok/s, TTFT 2.12 s (first request of a task), 1.21 s later
tok/s = output tokens per second per request, median. TTFT = p50.
Apps: nipale-ai.github.io/local-ai…
Oct 3, 2026 · 6:04 PM UTC
1
35
