Important caveat: this is a serving-efficiency experiment, not a quality benchmark. Thinking High consumed the full 1,000-token allowance as hidden reasoning and never reached a visible answer. A fair quality comparison needs a larger reasoning budget plus a separate visible-answer allowance.
Configuration: Apple M5 / 32 GB unified memory, LM Studio, Qwen 3.8 27B Q4_K_M, 32K context, four slots, MTP speculative decoding on, temperature 0, three workloads, two repeats per condition.