running PrismML's Bonsai 2 27B on a single RTX 3060. 12GB VRAM. (config below)
220K context, ~35 tok/s decode, ~550 tok/s prefill.
a 27B at 1.72 bits/weight, real ternary. two packings: PQ2_0 (7.21 GB) gets 220K and the numbers above, PTQ1_0 (5.95 GB) is smaller so it takes the full 262K but runs slower -- 26 tok/s decode, 260 prefill. max context go PTQ1_0, speed go PQ2_0.
needs PrismML's llama.cpp fork, kernels aren't upstream yet. stock llama.cpp rejects these files. day one config, expect it to move.
config:
llama-server
-m Ternary-Bonsai-2-27B-PQ2_0.gguf
-ngl 999 -c 220000 -fa on --jinja
-np 1
--cache-type-k q4_0 --cache-type-v q4_0
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
--presence-penalty 0.0 --repeat-penalty 1.0
--reasoning on
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.