Darkbloom is the first model provider to support Ternary Bonsai 2 27B -- concentrated intelligence that fits on your phone.
Try it now:
console.darkbloom.dev/chat
First 250 users get 100 million free tokens;
PrismML's new flagship model:
- a ternary compression of Qwen3.8 27B at 2bits per parameter.
- 8.5 GB total, 5x smaller than original
- keeps 98.2% of the Qwen's FP16 benchmark performance.
- 75% cheaper than Qwen 27B.
Qwen3.8 27B already operates comparably with Opus 4.6 and 5.6 Luna on certain tasks. This one does it in the memory of a phone. 262K context, image input, Apache 2.0.
From our first run on the network, on a single M5 Max with no caching:
- 35 tok/s decode at 1K context,
- 31 tok/s at 10K,
- 19 tok/s at 50K.
But we expect more performance gain coming in a few weeks! That's a full 27B reasoning model running comfortably on any Mac.
1,000+ Macs are serving on Darkbloom right now. Go try it out!!
Thank you to
@BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords and the whole PrismML team. This is exactly the kind of model Darkbloom was built for.
You can read the essay by Bonsai on Why Local AI Matters:
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.