Is this the Qwen3.8-27B fine-tune everyone has been asking for?
It has the two fixes people keep asking for in Qwen3.8-27B. (catchy name: ThinkingCap-Qwen3.8-27B-Uncensored-Heretic-GGUF)
Less overthinking + fewer refusals.
Stats ... ๐ง Qwen3.8-27B
โ ๐ ThinkingCap
โ ๐ Heretic
โ ๐ฆ GGUF
ThinkingCap reports...
โก 37.2% fewer thinking tokens
๐ 86.6 โ 85.8 avg accuracy
And an independent llama.cpp Aider test today found:
Vanilla Qwen:
๐ง 12,547 median tokens
โฑ๏ธ 1,481 sec/case
ThinkingCap:
๐ง 7,436 tokens
โฑ๏ธ 777 sec/case
โฆwith the SAME 77.6% retry-pass score in that test.
Then Heretic removes most refusal behavior.
Now mradermacher has local GGUFs:
๐ฆ IQ4_XS โ 15.3GB
๐ฅ Q4_K_M โ 16.6GB
๐ Q6_K โ 22.2GB
Runs on ...
โ
llama.cpp
โ
LM Studio
โ
Ollama
โ
multimodal
โ
local agents
So this is basically ๐ Qwen3.8-27B that wastes fewer tokens, refuses less, and fits on normal Local AI hardware. ๐
โ ๏ธ Important: ThinkingCap uses a PolyForm Small Business license, not Qwen's original Apache-2.0 license.
๐ Link in ALT
ALT https://huggingface.co/mradermacher/ThinkingCap-Qwen3.8-27B-Uncensored-Heretic-GGUF