Bonsai 2 27B
Bonsai 2 27B doesn't use the standard quantization ladder; it ships as fixed builds needing 7.2 GB at Ternary (Q2_0). 111 GPUs we track can run at least one build fully in VRAM at 8k context.
111 GPUs run this natively · 5 with CPU offload
- Ternary (Q2_0) total
- 7.2 GB
- at 8k context
- Smallest GPU
- 8 GB
- NVIDIA RTX 5060 Ti 8GB, at Ternary (Q2_0)
- KV cache, full context
- 17.2 GB
- 16 of 64 layers cache
- Inputs
- Text · Image
- Apache 2.0, released September 2026
- Overall retention
- 98.2%
- vs Qwen 3.8 27B FP16, PrismML's own suite
- Predecessor
- Bonsai 27B
- 94.6% retention vs Qwen 3.6 27B, Jul 2026
Bonsai 2 27B vs Qwen 3.8 27B: size & quality
Same weights as Qwen 3.8 27B, requantized to a fraction of the size.
Weights size
Quality retained
Quality retained is Qwen 3.8 27B's own FP16 output as the 100% reference, against PrismML's vendor-reported retention for each build, not yet verified by an independent leaderboard.
Bonsai 2 27B is a 27B parameter dense model developed by PrismML. Announced September 18, 2026 by PrismML, the second release in its Bonsai requantization line: a proprietary ternary (roughly 1.58 bits/weight) requantization of Alibaba's Qwen3.8-27B, identical underlying weights, not a new pretrain or finetune. Ships as a single fixed build under Apache 2.0, unlike the first Bonsai 27B's two-build lineup (1-bit and Ternary); PrismML's launch post frames this release as a quality-focused follow-up on the same ~5.9GB footprint as the original Ternary build, two months after that first release.
To run Bonsai 2 27B locally: At 5.9GB the single Ternary build is small enough for 8GB GPUs, base-tier Apple Silicon, or high-end phones, the same footprint class as the original Bonsai 27B's Ternary build despite the newer, larger Qwen3.8-27B base it's requantized from. PrismML's launch post doesn't publish tok/s figures for this release the way it did for the original Bonsai 27B, so no vendor-reported throughput numbers are available yet.
PrismML's own six-category capability table (Knowledge & Reasoning, Math, Coding, Agentic & Tool Calling, Instruction Following, Vision) reports 98.2% aggregate retention against the Qwen3.8-27B checkpoint it's requantized from, up from the original Bonsai 27B's 94.6% retention against Qwen3.6-27B. The gains concentrate in coding and agentic tool-calling, and Instruction Following actually exceeds the FP16 base model (82.66 vs 81.25, a 102% retention figure), the one category where ternary quantization apparently helps rather than hurts. Not yet listed on any independent leaderboard.
Available builds
Unlike models with a full FP32–Q2_K ladder, Bonsai 2 27B ships in 1 pre-baked low-bit builds only. The table below uses 8k of context as its baseline.
| Build | Bits/weight | Weights | KV cache | Total | Quality vs FP16 |
|---|---|---|---|---|---|
| Ternary (Q2_0) | 1.58 | 5.9 GB | 0.54 GB | 7.2 GB | ~98.2%* |
*These quality figures come from the vendor's own benchmarks and haven't been confirmed by an independent leaderboard yet. KV cache is shown at 8k context (FP16).
Benchmarks
Estimated, not measured: derived from Qwen 3.8 27B's real score times the best build's vendor-reported quality retention (see the chart above), not independently verified, and not a leaderboard result for Bonsai 2 27B itself.
PrismML's own capability table: what a 1.58-bit requantization actually costs
PrismML's launch announcement is unusually specific about where quality is lost, and, in one category, gained: the same six-category in-house capability suite run on the FP16 Qwen 3.8 27B base checkpoint and on this Ternary requantization of it.
PrismML (@PrismML on X), Bonsai 2 27B launch announcement, 18 September 2026. PrismML's own in-house capability suite; not yet listed on any independent leaderboard or independently reproduced.
Five of six categories retain 96.3%-99.5% of the FP16 base model's score, and the quantized build actually edges past it on Instruction Following (82.66 vs 81.25, a 102% retention figure), the one category where PrismML's own numbers show ternary quantization helping rather than hurting. Vision takes the largest hit (96.3% retention), unsurprising for a checkpoint compressed to 1.58 bits/weight. The reported 98.2% overall retention is a clear jump from the original Bonsai 27B's 94.6% retention against its own FP16 base two months earlier.
GPUs that run Bonsai 2 27B natively (111)
- NVIDIA RTX 5090Ternary (Q2_0) · 181 t/s
- NVIDIA RTX 5080Ternary (Q2_0) · 96.9 t/s
- NVIDIA RTX 5070 TiTernary (Q2_0) · 90.5 t/s
- NVIDIA RTX 5070Ternary (Q2_0) · 67.9 t/s
- NVIDIA RTX 5060 Ti 16GBTernary (Q2_0) · 45.2 t/s
Show 106 more
- NVIDIA RTX 5060 Ti 8GBTernary (Q2_0) · 45.2 t/s
- NVIDIA RTX 5060Ternary (Q2_0) · 45.2 t/s
- NVIDIA RTX 5050Ternary (Q2_0) · 32.3 t/s
- NVIDIA RTX 4090Ternary (Q2_0) · 101.8 t/s
- NVIDIA RTX 4080Ternary (Q2_0) · 72.4 t/s
- NVIDIA RTX 4070 Ti SUPERTernary (Q2_0) · 67.9 t/s
- NVIDIA RTX 4070 TiTernary (Q2_0) · 50.9 t/s
- NVIDIA RTX 4070 SUPERTernary (Q2_0) · 50.9 t/s
- NVIDIA RTX 4070Ternary (Q2_0) · 50.9 t/s
- NVIDIA RTX 4060 Ti 16GBTernary (Q2_0) · 29.1 t/s
- NVIDIA RTX 4060Ternary (Q2_0) · 27.5 t/s
- NVIDIA RTX 3090Ternary (Q2_0) · 94.5 t/s
- NVIDIA RTX 3090 TiTernary (Q2_0) · 101.8 t/s
- NVIDIA RTX 3080 10GBTernary (Q2_0) · 76.7 t/s
- NVIDIA RTX 3060 12GBTernary (Q2_0) · 36.4 t/s
- NVIDIA B300 288GBTernary (Q2_0) · 807.8 t/s
- NVIDIA B200 180GBTernary (Q2_0) · 807.8 t/s
- NVIDIA H200 141GBTernary (Q2_0) · 484.7 t/s
- NVIDIA H100 80GBTernary (Q2_0) · 338.3 t/s
- NVIDIA A100 80GBTernary (Q2_0) · 205.9 t/s
- NVIDIA A100 40GBTernary (Q2_0) · 157 t/s
- NVIDIA L40STernary (Q2_0) · 87.2 t/s
- NVIDIA RTX A6000Ternary (Q2_0) · 77.6 t/s
- NVIDIA RTX 4000 AdaTernary (Q2_0) · 32.3 t/s
- NVIDIA RTX 4500 AdaTernary (Q2_0) · 43.6 t/s
- NVIDIA RTX 5000 AdaTernary (Q2_0) · 58.2 t/s
- NVIDIA RTX 6000 AdaTernary (Q2_0) · 96.9 t/s
- NVIDIA RTX Pro 6000Ternary (Q2_0) · 135.7 t/s
- NVIDIA DGX Spark (128GB)Ternary (Q2_0) · 27.6 t/s
- AMD Radeon RX 7900 XTXTernary (Q2_0) · 96.9 t/s
- AMD Radeon RX 7900 XTTernary (Q2_0) · 80.8 t/s
- AMD Radeon RX 7900 GRETernary (Q2_0) · 58.2 t/s
- AMD Radeon RX 6800 XTTernary (Q2_0) · 51.7 t/s
- AMD Radeon PRO W7800Ternary (Q2_0) · 58.2 t/s
- AMD Radeon PRO W7900Ternary (Q2_0) · 87.2 t/s
- AMD Instinct MI300XTernary (Q2_0) · 535.2 t/s
- AMD Radeon AI PRO R9700 32GBTernary (Q2_0) · 64.6 t/s
- AMD Strix Halo (128GB)Ternary (Q2_0) · 25.9 t/s
- AMD Strix Halo (96GB)Ternary (Q2_0) · 25.9 t/s
- AMD Strix Halo (64GB)Ternary (Q2_0) · 25.9 t/s
- AMD Strix Halo (32GB)Ternary (Q2_0) · 25.9 t/s
- Apple M5 Ultra (512GB)Ternary (Q2_0) · 149.1 t/s
- Apple M5 Ultra (256GB)Ternary (Q2_0) · 149.1 t/s
- Apple M5 Ultra (96GB)Ternary (Q2_0) · 149.1 t/s
- Apple M5 Max (128GB)Ternary (Q2_0) · 76.3 t/s
- Apple M5 Max (64GB)Ternary (Q2_0) · 76.3 t/s
- Apple M5 Max (48GB)Ternary (Q2_0) · 76.3 t/s
- Apple M5 Max (36GB)Ternary (Q2_0) · 57.2 t/s
- Apple M5 Pro (64GB)Ternary (Q2_0) · 38.2 t/s
- Apple M5 Pro (48GB)Ternary (Q2_0) · 38.2 t/s
- Apple M5 Pro (24GB)Ternary (Q2_0) · 38.2 t/s
- Apple M5 (32GB)Ternary (Q2_0) · 19 t/s
- Apple M5 (16GB)Ternary (Q2_0) · 19 t/s
- Apple M6 (32GB)Ternary (Q2_0) · 21.1 t/s
- Apple M6 (16GB)Ternary (Q2_0) · 21.1 t/s
- Apple M4 Max (128GB)Ternary (Q2_0) · 67.9 t/s
- Apple M4 Max (64GB)Ternary (Q2_0) · 67.9 t/s
- Apple M4 Max (48GB)Ternary (Q2_0) · 67.9 t/s
- Apple M4 Max (36GB)Ternary (Q2_0) · 51 t/s
- Apple M4 Pro (48GB)Ternary (Q2_0) · 33.9 t/s
- Apple M4 Pro (24GB)Ternary (Q2_0) · 33.9 t/s
- Apple M4 (32GB)Ternary (Q2_0) · 14.9 t/s
- Apple M4 (16GB)Ternary (Q2_0) · 14.9 t/s
- Apple M3 Ultra (512GB)Ternary (Q2_0) · 101.8 t/s
- Apple M3 Ultra (256GB)Ternary (Q2_0) · 101.8 t/s
- Apple M3 Ultra (96GB)Ternary (Q2_0) · 101.8 t/s
- Apple M3 Max (128GB)Ternary (Q2_0) · 49.7 t/s
- Apple M3 Max (96GB)Ternary (Q2_0) · 37.3 t/s
- Apple M3 Max (64GB)Ternary (Q2_0) · 49.7 t/s
- Apple M3 Max (48GB)Ternary (Q2_0) · 49.7 t/s
- Apple M3 Max (36GB)Ternary (Q2_0) · 37.3 t/s
- Apple M3 Pro (36GB)Ternary (Q2_0) · 18.6 t/s
- Apple M3 Pro (18GB)Ternary (Q2_0) · 18.6 t/s
- Apple M3 (24GB)Ternary (Q2_0) · 12.4 t/s
- Apple M3 (16GB)Ternary (Q2_0) · 12.4 t/s
- Apple M2 Ultra (192GB)Ternary (Q2_0) · 99.4 t/s
- Apple M2 Ultra (64GB)Ternary (Q2_0) · 99.4 t/s
- Apple M2 Max (96GB)Ternary (Q2_0) · 49.7 t/s
- Apple M2 Max (64GB)Ternary (Q2_0) · 49.7 t/s
- Apple M2 Max (32GB)Ternary (Q2_0) · 49.7 t/s
- Apple M2 Pro (32GB)Ternary (Q2_0) · 24.9 t/s
- Apple M2 Pro (16GB)Ternary (Q2_0) · 24.9 t/s
- Apple M2 (24GB)Ternary (Q2_0) · 12.4 t/s
- Apple M2 (16GB)Ternary (Q2_0) · 12.4 t/s
- Apple M1 Ultra (128GB)Ternary (Q2_0) · 99.4 t/s
- Apple M1 Ultra (64GB)Ternary (Q2_0) · 99.4 t/s
- Apple M1 Max (64GB)Ternary (Q2_0) · 49.7 t/s
- Apple M1 Max (32GB)Ternary (Q2_0) · 49.7 t/s
- Apple M1 Pro (32GB)Ternary (Q2_0) · 24.9 t/s
- Apple M1 Pro (16GB)Ternary (Q2_0) · 24.9 t/s
- Apple M1 (16GB)Ternary (Q2_0) · 8.5 t/s
- Intel Arc B580 12GBTernary (Q2_0) · 46 t/s
- Intel Arc B570 10GBTernary (Q2_0) · 38.4 t/s
- Intel Arc Pro B70 32GBTernary (Q2_0) · 61.4 t/s
- Intel Arc Pro B60 24GBTernary (Q2_0) · 38.4 t/s
- Intel Arc Pro B50 16GBTernary (Q2_0) · 22.6 t/s
- Intel Arc A770 16GBTernary (Q2_0) · 56.5 t/s
- Intel Arc A770 8GBTernary (Q2_0) · 51.7 t/s
- Intel Arc A750 8GBTernary (Q2_0) · 51.7 t/s
- Intel Arc A580 8GBTernary (Q2_0) · 51.7 t/s
- Intel Arc Pro A60 12GBTernary (Q2_0) · 38.8 t/s
- Intel Data Center GPU Max 1550Ternary (Q2_0) · 330.8 t/s
- Intel Data Center GPU Max 1100Ternary (Q2_0) · 124.1 t/s
- Intel Arc 140V (32GB)Ternary (Q2_0) · 13.8 t/s
- Intel Arc 140V (16GB)Ternary (Q2_0) · 13.8 t/s
- Intel Arc 130V (16GB)Ternary (Q2_0) · 13.8 t/s
Plus 5 GPUs that run it with CPU offload (slower)
- Intel Arc A380 6GBTernary (Q2_0) · 13.1 t/s
- Intel Arc A310 4GBTernary (Q2_0) · 6.6 t/s
- Intel Arc Pro A50 6GBTernary (Q2_0) · 13.3 t/s
- Intel Arc Pro A40 6GBTernary (Q2_0) · 13.3 t/s
- CPU only (system RAM)Ternary (Q2_0) · 6.2 t/s
Notes
Third-party requantization of Qwen3.8-27B by PrismML, same weights, not a retrain, and the second release in PrismML's Bonsai line (the first, requantized from Qwen3.6-27B, shipped July 2026). Doesn't use the standard FP32-Q2_K ladder shown for other models; it ships as a single fixed sub-2-bit build (see the quantization table on this page). PrismML's own announcement (X/Twitter @PrismML, 18 September 2026) reports a six-category capability table against Qwen 3.6 27B and Qwen 3.8 27B, with an aggregate 98.2% of Qwen3.8-27B's benchmark performance retained at this footprint, PrismML's own in-house suite, not yet independently verified by a third party. The Terminal-Bench score shown is CanItRun's own estimate (Qwen3.8-27B's score x the Agentic & Tool Calling retention PrismML reported), not a number PrismML or any leaderboard has published for this model directly.
Compare Bonsai 2 27B with other models
So should you run it over the original Bonsai 27B?
If you already run the original Bonsai 27B and have the disk space, yes. The download totals 5.9 GB of weights (7.2 GB once KV cache and overhead are added at 8k context), the same footprint class as the original's own Ternary build despite being requantized from the newer, stronger Qwen 3.8 27B, and PrismML's own numbers show it holding onto more of its base model's quality (98.2% vs 94.6% retention) than the original did. It's also roughly 9.2x smaller than the 54.0 GB of BF16 weights Qwen 3.8 27B itself ships as, close to PrismML's own rounded '9x' framing. What it isn't is a free way to get frontier 27B-class reasoning: 1.58 bits/weight is still an aggressive compression, and none of PrismML's benchmark claims have independent verification yet. For anyone not already committed to the Bonsai line's ultra-low-bit format, a standard Q4_K_M build of Qwen 3.8 27B itself (about 19.0 GB at 8k context) is the safer choice for quality-sensitive work; reach for this release only when an 8GB-class GPU, base-tier Apple Silicon, or a phone is the hard constraint.
Frequently asked questions
- What are the VRAM requirements for Bonsai 2 27B?
- Bonsai 2 27B doesn't use the standard quantization ladder; it ships as 1 fixed builds: Ternary (Q2_0) (7.2 GB). These figures assume 8k of context; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Bonsai 2 27B have?
- Bonsai 2 27B has 27 billion parameters, the same count as the base model it's derived from, since this release is a post-training requantization rather than a retrain.
- What quantization levels does Bonsai 2 27B come in?
- Bonsai 2 27B skips the usual FP16-to-Q2_K ladder entirely. It's only available as: Ternary (Q2_0) at 1.58 bits/weight (7.2 GB total, ~98.2% of FP16 quality per the vendor).
- Can Bonsai 2 27B run on a small GPU?
- Yes. The Ternary (Q2_0) build needs only 7.2 GB, which fits on most 8 GB+ GPUs, entry-level Apple Silicon, and even high-end phones per the vendor's own claims.
- What GPU do I need to run Bonsai 2 27B locally?
- Every build is small enough for almost any GPU released in the last several years. The Ternary (Q2_0) build (7.2 GB) needs a bit more headroom for better quality, but all of them fit comfortably on an 8 GB card or larger.
- How is Ternary Bonsai 2 27B different from the original Bonsai 27B?
- Both are proprietary low-bit requantizations from PrismML, not retrains, but of different base models: the original Bonsai 27B (July 2026) requantizes Qwen 3.6 27B and ships two builds (1-bit and Ternary); Bonsai 2 27B (September 2026) requantizes the newer, stronger Qwen 3.8 27B and ships a single Ternary build. PrismML's own numbers show the newer release retaining more of its base model's quality, 98.2% versus the original Ternary build's 94.6%, at essentially the same 5.9-7.2 GB footprint.
- Can I run Ternary Bonsai 2 27B on an 8GB GPU?
- It's built for exactly that class of hardware: 8GB GPUs, base-tier Apple Silicon, or high-end phones, the same class the original Bonsai 27B targets. At a full 262,144-token context the KV cache grows further, so check this page's per-GPU table before assuming a specific 8GB card clears it at long context, not just at the 8k baseline.
- Why does Instruction Following score higher on the quantized build than on the FP16 base model?
- PrismML's own table reports it that way (82.66 vs Qwen 3.8 27B's 81.25, a 102% retention figure), and it's the only one of the six categories where that happens. PrismML doesn't explain the effect, and a gap that small on an unpublished-size eval set is also consistent with ordinary run-to-run noise rather than quantization genuinely improving instruction-following. Treat it as a real reported number, not yet a confirmed mechanism.