Bonsai 2 27B vs Bonsai 27B
Side-by-side VRAM requirements, benchmark scores, and GPU compatibility for local AI inference.
Quick verdict
Bonsai 2 27B is more hardware-efficient: it needs 7.2 GB at its Ternary (Q2_0) build vs 8.7 GB for Bonsai 27B's Ternary (Q2_0), fitting on 111 GPUs natively.
Analysis
Bonsai 2 27B and the original Bonsai 27B are both proprietary low-bit requantizations from PrismML, same idea, different base checkpoint. The original (July 2026) requantizes Qwen 3.6 27B; the sequel (September 2026) requantizes the newer, stronger Qwen 3.8 27B. Since both base models share the same real hybrid architecture (16 of 64 layers keep a KV cache; the other 48 run linear attention with a fixed-size state), the KV-cache math is identical between the two Bonsai releases; the interesting differences are in what PrismML actually shipped and how much quality each retained.
The original Bonsai 27B ships two fixed builds: a 1-bit (Q1_0) file at 3.8 GB weights retaining 89.5% of FP16 quality, and a Ternary (Q2_0) file at 7.2 GB retaining 94.6%. Bonsai 2 27B ships only one: a Ternary (Q2_0) file at 5.9 GB, smaller than the original's own Ternary build despite the identical 27B parameter count and the same 1.58 bits/weight, so PrismML's packing improved on top of retargeting the base model. It also retains more: PrismML's own six-category capability table for Bonsai 2 27B reports 98.2% aggregate retention against Qwen 3.8 27B, against the original's 94.6% retention against Qwen 3.6 27B, and Bonsai 2 27B is the rare case where one category (Instruction Following) actually scores above its FP16 base rather than below it. Neither release has independent third-party benchmark verification yet; both are PrismML's own reported figures.
Bottom line: For anyone choosing between the two today, Bonsai 2 27B is the better pick: a smaller download than the original's Ternary build, higher reported quality retention, and a newer, stronger base model underneath. The only reason to reach for the original instead is its 1-bit build, at 3.8 GB the smallest Bonsai file that exists; Bonsai 2 27B doesn't offer an equivalent sub-4GB option, so a 6-8GB device that can't fit even 5.9 GB is the one case where the original's extra build still matters.
VRAM at each real build (8k context)
Bonsai 2 27B and Bonsai 27B both ship as a handful of fixed prebuilt files rather than a standard quantization ladder, so the two don't share equivalent bit-widths to line up row for row. Ranked smallest to largest instead.
| Rank | Bonsai 2 27B | GB | Bonsai 27B | GB |
|---|---|---|---|---|
| 1 | Ternary (Q2_0) | 7.2 GB | 1-bit (Q1_0) | 4.9 GB |
| 2 | — | — | Ternary (Q2_0) | 8.7 GB |
Each column is that model's own real builds, smallest first; rank pairs them by position, not by matching precision.
Model specifications
| Spec | Bonsai 2 27B | Bonsai 27B |
|---|---|---|
| Org | PrismML | PrismML |
| Parameters | 27B | 27B |
| Architecture | Dense | Dense |
| Context | 256k tokens | 256k tokens |
| Modalities | text, vision | text, vision |
| License | Apache 2.0 | Apache 2.0 |
| Commercial | Yes | Yes |
| Released | 2026-09-18 | 2026-07-14 |
| GPUs (native) | 111 / 119 | 114 / 119 |
Benchmark scores
| Benchmark | Bonsai 2 27B | Bonsai 27B |
|---|---|---|
| Terminal-Bench 2.1 | ~71.0 | N/A |
Green = higher score (better). N/A = not yet available. ~ = inherited from a base model, not independently reported for that release itself.
GPUs that run only Bonsai 2 27B(0)
Every GPU that runs Bonsai 2 27B also runs Bonsai 27B.
GPUs that run only Bonsai 27B(3)
GPUs that run both natively(111)
- NVIDIA RTX 509032 GB
- NVIDIA RTX 508016 GB
- NVIDIA RTX 5070 Ti16 GB
- NVIDIA RTX 507012 GB
- NVIDIA RTX 5060 Ti 16GB16 GB
- NVIDIA RTX 5060 Ti 8GB8 GB
- NVIDIA RTX 50608 GB
- NVIDIA RTX 50508 GB
- NVIDIA RTX 409024 GB
- NVIDIA RTX 408016 GB
- NVIDIA RTX 4070 Ti SUPER16 GB
- NVIDIA RTX 4070 Ti12 GB
- +99 more GPUs run both
Which should you use?
- • It's the newer release (2026-09-18 vs 2026-07-14); check the benchmark table above for what actually improved
- No clear spec advantage over Bonsai 2 27B, see the benchmark and VRAM tables above.
Frequently asked questions
- Which is better, Bonsai 2 27B or Bonsai 27B?
- Bonsai 2 27B is more hardware-efficient, needing 7.2 GB at its Ternary (Q2_0) build vs 8.7 GB for Bonsai 27B's Ternary (Q2_0). Bonsai 27B runs on more GPUs natively (114 vs 111).
- How much VRAM does Bonsai 2 27B need vs Bonsai 27B?
- At 8k context, Bonsai 2 27B needs approximately 7.2 GB of VRAM at its Ternary (Q2_0) build, while Bonsai 27B needs 8.7 GB at its Ternary (Q2_0) build. At the largest build each ships, Bonsai 2 27B requires 7.2 GB (Ternary (Q2_0)) vs 8.7 GB (Ternary (Q2_0)) for Bonsai 27B.
- Can you run Bonsai 2 27B on the same GPUs as Bonsai 27B?
- Yes, 111 GPUs can run both natively in VRAM, including NVIDIA RTX 5090, NVIDIA RTX 5080, NVIDIA RTX 5070 Ti. However, no GPU can run Bonsai 2 27B without also fitting Bonsai 27B, and 3 GPUs can run Bonsai 27B but not Bonsai 2 27B.
- What is the difference between Bonsai 2 27B and Bonsai 27B?
- Bonsai 2 27B has 27B parameters (dense) with a 256k context window. Bonsai 27B has 27B parameters (dense) with a 256k context window.
- Which model fits in 24 GB of VRAM, Bonsai 2 27B or Bonsai 27B?
- Both fit in 24 GB of VRAM at their respective recommended builds: Bonsai 2 27B (Ternary (Q2_0)) needs 7.2 GB and Bonsai 27B (Ternary (Q2_0)) needs 8.7 GB.