Bonsai 2 27B

Bonsai 2 27B doesn't use the standard quantization ladder; it ships as fixed builds needing 7.2 GB at Ternary (Q2_0). 111 GPUs we track can run at least one build fully in VRAM at 8k context.

111 GPUs run this natively · 5 with CPU offload

PrismML27B params256k contextApache 2.0Commercial use ok
Ternary (Q2_0) total
7.2 GB
at 8k context
Smallest GPU
8 GB
NVIDIA RTX 5060 Ti 8GB, at Ternary (Q2_0)
KV cache, full context
17.2 GB
16 of 64 layers cache
Inputs
Text · Image
Apache 2.0, released September 2026
Overall retention
98.2%
vs Qwen 3.8 27B FP16, PrismML's own suite
Predecessor
Bonsai 27B
94.6% retention vs Qwen 3.6 27B, Jul 2026

Bonsai 2 27B vs Qwen 3.8 27B: size & quality

Same weights as Qwen 3.8 27B, requantized to a fraction of the size.

Weights size

Qwen 3.8 27B FP16
54.0 GB
Qwen 3.8 27B Q4_K_M
16.4 GB
Bonsai Ternary
5.9 GB

Quality retained

Qwen 3.8 27B FP16
100% (reference)
Bonsai Ternary
~98.2%

Quality retained is Qwen 3.8 27B's own FP16 output as the 100% reference, against PrismML's vendor-reported retention for each build, not yet verified by an independent leaderboard.

Bonsai 2 27B is a 27B parameter dense model developed by PrismML. Announced September 18, 2026 by PrismML, the second release in its Bonsai requantization line: a proprietary ternary (roughly 1.58 bits/weight) requantization of Alibaba's Qwen3.8-27B, identical underlying weights, not a new pretrain or finetune. Ships as a single fixed build under Apache 2.0, unlike the first Bonsai 27B's two-build lineup (1-bit and Ternary); PrismML's launch post frames this release as a quality-focused follow-up on the same ~5.9GB footprint as the original Ternary build, two months after that first release.

To run Bonsai 2 27B locally: At 5.9GB the single Ternary build is small enough for 8GB GPUs, base-tier Apple Silicon, or high-end phones, the same footprint class as the original Bonsai 27B's Ternary build despite the newer, larger Qwen3.8-27B base it's requantized from. PrismML's launch post doesn't publish tok/s figures for this release the way it did for the original Bonsai 27B, so no vendor-reported throughput numbers are available yet.

PrismML's own six-category capability table (Knowledge & Reasoning, Math, Coding, Agentic & Tool Calling, Instruction Following, Vision) reports 98.2% aggregate retention against the Qwen3.8-27B checkpoint it's requantized from, up from the original Bonsai 27B's 94.6% retention against Qwen3.6-27B. The gains concentrate in coding and agentic tool-calling, and Instruction Following actually exceeds the FP16 base model (82.66 vs 81.25, a 102% retention figure), the one category where ternary quantization apparently helps rather than hurts. Not yet listed on any independent leaderboard.

Available builds

Unlike models with a full FP32–Q2_K ladder, Bonsai 2 27B ships in 1 pre-baked low-bit builds only. The table below uses 8k of context as its baseline.

BuildBits/weightWeightsKV cacheTotalQuality vs FP16
Ternary (Q2_0)1.585.9 GB0.54 GB7.2 GB~98.2%*

*These quality figures come from the vendor's own benchmarks and haven't been confirmed by an independent leaderboard yet. KV cache is shown at 8k context (FP16).

Benchmarks

Estimated, not measured: derived from Qwen 3.8 27B's real score times the best build's vendor-reported quality retention (see the chart above), not independently verified, and not a leaderboard result for Bonsai 2 27B itself.

PrismML's own capability table: what a 1.58-bit requantization actually costs

PrismML's launch announcement is unusually specific about where quality is lost, and, in one category, gained: the same six-category in-house capability suite run on the FP16 Qwen 3.8 27B base checkpoint and on this Ternary requantization of it.

Knowledge & Reasoning
Bonsai 2 27B
84.0
Qwen 3.8 27B base
86.7
Math
Bonsai 2 27B
96.6
Qwen 3.8 27B base
97.1
Coding
Bonsai 2 27B
81.6
Qwen 3.8 27B base
82.2
Agentic & Tool Calling
Bonsai 2 27B
77.6
Qwen 3.8 27B base
79.7
Instruction Following
Bonsai 2 27B
82.7
Qwen 3.8 27B base
81.3
Vision
Bonsai 2 27B
78.6
Qwen 3.8 27B base
81.6
Overall
Bonsai 2 27B
83.9
Qwen 3.8 27B base
85.4

PrismML (@PrismML on X), Bonsai 2 27B launch announcement, 18 September 2026. PrismML's own in-house capability suite; not yet listed on any independent leaderboard or independently reproduced.

Five of six categories retain 96.3%-99.5% of the FP16 base model's score, and the quantized build actually edges past it on Instruction Following (82.66 vs 81.25, a 102% retention figure), the one category where PrismML's own numbers show ternary quantization helping rather than hurting. Vision takes the largest hit (96.3% retention), unsurprising for a checkpoint compressed to 1.58 bits/weight. The reported 98.2% overall retention is a clear jump from the original Bonsai 27B's 94.6% retention against its own FP16 base two months earlier.

GPUs that run Bonsai 2 27B natively (111)

Show 106 more
Plus 5 GPUs that run it with CPU offload (slower)

Notes

Third-party requantization of Qwen3.8-27B by PrismML, same weights, not a retrain, and the second release in PrismML's Bonsai line (the first, requantized from Qwen3.6-27B, shipped July 2026). Doesn't use the standard FP32-Q2_K ladder shown for other models; it ships as a single fixed sub-2-bit build (see the quantization table on this page). PrismML's own announcement (X/Twitter @PrismML, 18 September 2026) reports a six-category capability table against Qwen 3.6 27B and Qwen 3.8 27B, with an aggregate 98.2% of Qwen3.8-27B's benchmark performance retained at this footprint, PrismML's own in-house suite, not yet independently verified by a third party. The Terminal-Bench score shown is CanItRun's own estimate (Qwen3.8-27B's score x the Agentic & Tool Calling retention PrismML reported), not a number PrismML or any leaderboard has published for this model directly.

Hugging Face ↗Released 2026-09-18

Compare Bonsai 2 27B with other models

So should you run it over the original Bonsai 27B?

If you already run the original Bonsai 27B and have the disk space, yes. The download totals 5.9 GB of weights (7.2 GB once KV cache and overhead are added at 8k context), the same footprint class as the original's own Ternary build despite being requantized from the newer, stronger Qwen 3.8 27B, and PrismML's own numbers show it holding onto more of its base model's quality (98.2% vs 94.6% retention) than the original did. It's also roughly 9.2x smaller than the 54.0 GB of BF16 weights Qwen 3.8 27B itself ships as, close to PrismML's own rounded '9x' framing. What it isn't is a free way to get frontier 27B-class reasoning: 1.58 bits/weight is still an aggressive compression, and none of PrismML's benchmark claims have independent verification yet. For anyone not already committed to the Bonsai line's ultra-low-bit format, a standard Q4_K_M build of Qwen 3.8 27B itself (about 19.0 GB at 8k context) is the safer choice for quality-sensitive work; reach for this release only when an 8GB-class GPU, base-tier Apple Silicon, or a phone is the hard constraint.

Frequently asked questions

What are the VRAM requirements for Bonsai 2 27B?
Bonsai 2 27B doesn't use the standard quantization ladder; it ships as 1 fixed builds: Ternary (Q2_0) (7.2 GB). These figures assume 8k of context; VRAM scales linearly with context length due to the KV cache.
How many parameters does Bonsai 2 27B have?
Bonsai 2 27B has 27 billion parameters, the same count as the base model it's derived from, since this release is a post-training requantization rather than a retrain.
What quantization levels does Bonsai 2 27B come in?
Bonsai 2 27B skips the usual FP16-to-Q2_K ladder entirely. It's only available as: Ternary (Q2_0) at 1.58 bits/weight (7.2 GB total, ~98.2% of FP16 quality per the vendor).
Can Bonsai 2 27B run on a small GPU?
Yes. The Ternary (Q2_0) build needs only 7.2 GB, which fits on most 8 GB+ GPUs, entry-level Apple Silicon, and even high-end phones per the vendor's own claims.
What GPU do I need to run Bonsai 2 27B locally?
Every build is small enough for almost any GPU released in the last several years. The Ternary (Q2_0) build (7.2 GB) needs a bit more headroom for better quality, but all of them fit comfortably on an 8 GB card or larger.
How is Ternary Bonsai 2 27B different from the original Bonsai 27B?
Both are proprietary low-bit requantizations from PrismML, not retrains, but of different base models: the original Bonsai 27B (July 2026) requantizes Qwen 3.6 27B and ships two builds (1-bit and Ternary); Bonsai 2 27B (September 2026) requantizes the newer, stronger Qwen 3.8 27B and ships a single Ternary build. PrismML's own numbers show the newer release retaining more of its base model's quality, 98.2% versus the original Ternary build's 94.6%, at essentially the same 5.9-7.2 GB footprint.
Can I run Ternary Bonsai 2 27B on an 8GB GPU?
It's built for exactly that class of hardware: 8GB GPUs, base-tier Apple Silicon, or high-end phones, the same class the original Bonsai 27B targets. At a full 262,144-token context the KV cache grows further, so check this page's per-GPU table before assuming a specific 8GB card clears it at long context, not just at the 8k baseline.
Why does Instruction Following score higher on the quantized build than on the FP16 base model?
PrismML's own table reports it that way (82.66 vs Qwen 3.8 27B's 81.25, a 102% retention figure), and it's the only one of the six categories where that happens. PrismML doesn't explain the effect, and a gap that small on an unpublished-size eval set is also consistent with ordinary run-to-run noise rather than quantization genuinely improving instruction-following. Treat it as a real reported number, not yet a confirmed mechanism.