🚨ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
📄 Blogpost: bottlecapai.com/post/thinkin…
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🔗 GGUF: huggingface.co/bottlecapai/T…
🚨 Same model size. Same GPU. ~4.7x more work done!
ThinkingCap: Qwen 3.8 reaches answers with about half the reasoning. On a shared server, that means 2x more users fit in memory at once, and each finishes 2x faster.
Running it solo? ~2x faster answers.
Serving a team? ~4.7x throughput. 🚀
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🚨 Same model size. Same GPU. ~4.7x more work done!
ThinkingCap: Qwen 3.8 reaches answers with about half the reasoning. On a shared server, that means 2x more users fit in memory at once, and each finishes 2x faster.
Running it solo? ~2x faster answers.
Serving a team? ~4.7x throughput. 🚀
🤗 Hugging Face: huggingface.co/bottlecapai/T…
ThinkingCap: Qwen3.8-27B now runs on a Mac and fits a 32GB machine!
4-bit MLX build - 21 GB, down from 52 GB
Performance is very close to the full precision version.
Vision capabilities and MTP speculative decoding are kept.
huggingface.co/bottlecapai/T…
🚨ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
📄 Blogpost: bottlecapai.com/post/thinkin…
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🔗 GGUF: huggingface.co/bottlecapai/T…
Actually we were very busy with another big project in works, but I am glad we finally got some time to polish Qwen 3.8 27b with our ThinkingCap. A lot positive feedback so far!🙏
🚨ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
📄 Blogpost: bottlecapai.com/post/thinkin…
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🔗 GGUF: huggingface.co/bottlecapai/T…
🚨ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
📄 Blogpost: bottlecapai.com/post/thinkin…
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🔗 GGUF: huggingface.co/bottlecapai/T…
🚨ThinkingCap: Qwen3.8-27B is out!
The most token-efficient Qwen3.8 27B available, this time optimized for agentic use.
🚀 Up to 65% fewer thinking tokens, 37% on average
🎯 Minimal accuracy loss
📈 Long-context retrieval actually went up 2.3pp
📄 Blogpost: bottlecapai.com/post/thinkin…
🤗 Hugging Face: huggingface.co/bottlecapai/T…
🔗 GGUF: huggingface.co/bottlecapai/T…
🏢 Enterprise use This release is trained at xhigh effort. For enterprise use, we are ready to fine-tune medium and low effort as well. Contact us at enterprise [at] bottlecapai.com