Bonsai 27B

Bonsai 27B doesn't use the standard quantization ladder; it ships as fixed builds needing 6.1 GB at 1-bit (Q1_0) or 9.9 GB at Ternary (Q2_0). 111 GPUs we track can run at least one build fully in VRAM at 8k context.

111 GPUs run this natively · 5 with CPU offload

PrismML27B params256k contextApache 2.0Commercial use ok

Bonsai 27B vs Qwen 3.6 27B: size & quality

Same weights as Qwen 3.6 27B, requantized to a fraction of the size.

Weights size

Qwen 3.6 27B FP16
54.0 GB
Qwen 3.6 27B Q4_K_M
16.4 GB
Bonsai 1-bit
3.8 GB
Bonsai Ternary
7.2 GB

Quality retained

Qwen 3.6 27B FP16
100% (reference)
Bonsai 1-bit
~89.5%
Bonsai Ternary
~94.6%

Quality retained is Qwen 3.6 27B's own FP16 output as the 100% reference, against PrismML's vendor-reported retention for each build, not yet verified by an independent leaderboard.

Bonsai 27B is a 27B parameter dense model developed by PrismML. Released July 2026 by PrismML as a proprietary low-bit requantization of Alibaba's Qwen3.6-27B: identical underlying weights, not a new pretrain or finetune. Two fixed builds, both Apache 2.0 (see the size chart above).

To run Bonsai 27B locally: Small enough for 8GB GPUs, base-tier Apple Silicon, or high-end phones. PrismML reports ~163 tok/s (1-bit) and ~134 tok/s (Ternary) on an RTX 5090, and ~87/58 tok/s on an M5 Max, vendor-reported, not independently verified.

Quality losses concentrate in agentic tool-calling and instruction-following rather than math (see the quality chart above). Not yet listed on any independent leaderboard.

Available builds

Bonsai 27B doesn't offer the standard quantization ladder; instead there are 2 fixed builds to choose from. Sizes assume 8k context throughout.

BuildBits/weightWeightsKV cacheTotalQuality vs FP16
1-bit (Q1_0)1.1253.8 GB1.61 GB6.1 GB~89.5%*
Ternary (Q2_0)1.587.2 GB1.61 GB9.9 GB~94.6%*

*These quality figures come from the vendor's own benchmarks and haven't been confirmed by an independent leaderboard yet. KV cache is shown at 8k context (FP16).

Benchmarks

Estimated, not measured: derived from Qwen 3.6 27B's real score times the best build's vendor-reported quality retention (see the chart above), not independently verified, and not a leaderboard result for Bonsai 27B itself.

GPUs that run Bonsai 27B natively (111)

Show 106 more
Plus 5 GPUs that run it with CPU offload (slower)

Notes

Third-party requantization of Qwen3.6-27B by PrismML, same weights, not a retrain. Doesn't use the standard FP32-Q2_K ladder shown for other models; it ships only as two fixed sub-2-bit builds (see the quantization table on this page). PrismML's own benchmark suite reports 89.5%–94.6% of FP16 quality retained across the two builds; not yet independently verified by a third party. The MMLU-Pro score shown is CanItRun's own estimate (Qwen3.6-27B's score x the Ternary build's retention), not a number PrismML or any leaderboard has published for Bonsai directly.

Hugging Face ↗Released 2026-07-14

Frequently asked questions

What are the VRAM requirements for Bonsai 27B?
Bonsai 27B doesn't use the standard quantization ladder; it ships as 2 fixed builds: 1-bit (Q1_0) (6.1 GB) and Ternary (Q2_0) (9.9 GB). These figures assume 8k of context; VRAM scales linearly with context length due to the KV cache.
How many parameters does Bonsai 27B have?
Bonsai 27B has 27 billion parameters, the same count as the base model it's derived from, since this release is a post-training requantization rather than a retrain.
What quantization levels does Bonsai 27B come in?
Bonsai 27B skips the usual FP16-to-Q2_K ladder entirely. It's only available as: 1-bit (Q1_0) at 1.125 bits/weight (6.1 GB total, ~89.5% of FP16 quality per the vendor); Ternary (Q2_0) at 1.58 bits/weight (9.9 GB total, ~94.6% of FP16 quality per the vendor).
Can Bonsai 27B run on a small GPU?
Yes. The 1-bit (Q1_0) build needs only 6.1 GB, which fits on most 8 GB+ GPUs, entry-level Apple Silicon, and even high-end phones per the vendor's own claims.
What GPU do I need to run Bonsai 27B locally?
Either build is small enough for almost any GPU released in the last several years. The Ternary (Q2_0) build (9.9 GB) needs a bit more headroom for better quality, but all of them fit comfortably on an 8 GB card or larger.