Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

At a Glance
| Item | Value |
|---|---|
| Repository | prism-ml/Ternary-Bonsai-2-27B-gguf-dev |
| Published | 2026-09-17 |
| License | apache-2.0 |
| Formats | GGUF |
| Source type | Unverified (not confirmed by a primary source) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
It is reported that Prism ML has released a development and testing GGUF build, prism-ml/Ternary-Bonsai-2-27B-gguf-dev, for “Bonsai 2 27B", a ternary weight model based on Qwen3.8-27B.
However, this news has not been officially confirmed and is currently considered unconfirmed information. It is reported that this build was prepared for the purpose of kernel development and testing for Q2_0 packing in llama.cpp, as well as verification for future upstream integration.
Specifications
- Parameters: 27B (base model)
- Architecture: Qwen3_5ForConditionalGeneration (base model)
- Context length: 262,144 tokens (base model)
Performance
Regarding the performance of the base model “Qwen3.8-27B", the comparison table provided in the model card from the publisher is as follows. Note that the comparison targets have been narrowed down to four.
Text Performance
→ Scroll horizontally to see all columns
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Opus4.6 Max | |
|---|---|---|---|---|
| Coding | ||||
| Agentic terminal coding Terminal Bench 2.1 (Terminus) | 73.0 | 63.4 | 64.0 | 78.2 |
| Agentic coding SWE-bench Pro | 61.7 | 53.5 | 57.6 | 53.4 |
| Repo-level code generation NL2Repo-Bench | 42.3 | 36.2 | 41.1 | 47.6 |
| Agentic coding DeepSWE 1.1 | 42.2 | 13.3 | 14.2 | — |
| Software engineering QwenSWEBench | 79.0 | 49.3 | 59.2 | 63.8 |
| Agent | ||||
| Long-horizon office work CoWorkBench | 70.7 | 61.0 | 65.1 | 68.2 |
| Professional job tasks JobBench | 33.4 | 21.8 | 27.6 | — |
| Frontier agentic tasks Agents’ Last Exam | Pass@1 20.4 Score 42.9 | Pass@1 10.6 Score 27.3 | Pass@1 13.2 Score 33.6 | — |
| General | ||||
| Instruction following IFBench | 79.5 | 69.1 | 79.1 | 62.5 |
| Scientific reasoning GPQA Diamond | 89.2 | 87.8 | 90.3 | 91.3 |
| Multidisciplinary reasoning HLE | 30.8 | 24.0 | 34.7 | 40.0 |
| Competitive coding LiveCodeBench v6 | 90.3 | 83.9 | 89.6 | 88.8 |
VL Performance
→ Scroll horizontally to see all columns
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Opus4.6 Max | |
|---|---|---|---|---|
| Agentic Multimodal Intelligence | ||||
| Computer use OSWorld-Verified | 84.3 | 63.9 | 73.3 | 72.7 |
| Browser use WebArena-Verified | 64.8 | 48.8 | 55.3 | — |
| Mobile use AndroidWorld | 81.9 | 70.3 | 81.0 | 62.0 |
| Application recreation RecreationBench | 47.1 | 29.8 | 30.2 | — |
| Multimodal tool use ClawEval-MM | Pass@3 57.4 Average 56.9 | Pass@3 42.6 Average 50.4 | Pass@3 57.4 Average 60.1 | Pass@3 52.5 Average 54.7 |
| Multimodal software engineering SWE-MM | 38.6 | 25.7 | 30.0 | 27.1 |
| Visual web development Vision2Web | 62.9 | 45.0 | 42.1 | — |
| General Multimodal Intelligence | ||||
| Visual math problem solving MathVision | Without CI 90.0 With CI 94.6 | Without CI 85.1 | Without CI 90.3 | Without CI 65.5 |
| General visual reasoning BabyVision | Without CI 65.7 With CI 85.6 | Without CI 28.9 | Without CI 64.7 With CI 70.4 | Without CI 12.6 |
| Scientific chart analysis CharXiv (RQ) | Without CI 83.7 With CI 90.2 | Without CI 78.4 | Without CI 85.8 With CI 85.9 | Without CI 66.0 |
| Document intelligence OmniDocBench 1.5 | 91.1 | 89.4 | 91.4 | 86.6 |
| Real-world perception RealWorldQA | 85.9 | 84.1 | 86.9 | 73.9 |
| Embodied intelligence ERQA | 65.5 | 62.5 | 69.8 | 40.8 |
According to the measurement results published by the base model’s creators, Qwen3.8-27B is reported to show high performance particularly in coding and agent tasks. It is reported to have recorded 61.7 on “SWE-bench Pro" (which measures the ability to resolve real GitHub issues) and 73.0 on “Terminal-Bench" (which measures terminal operation tasks), demonstrating capabilities that surpass “Opus4.6 Max" and “Qwen3.7-Plus" in some items. On the other hand, it is reported to fall behind in some extremely difficult tasks, remaining at 30.8 on “HLE" (which handles expert-level extremely difficult questions) compared to the top-tier model “Opus4.6 Max" at 40.0. Additionally, it scored 89.2 on “GPQA Diamond" (which measures scientific difficult questions), demonstrating high scientific reasoning capabilities approaching upper-tier models.
Strengths and Use Cases
The base model “Qwen3.8-27B" is reported to excel in coding, mathematics, long-context reading, multilingual support, agent tasks, and multimodal visual understanding (Vision-Language) including images and videos.
However, the newly released “Ternary-Bonsai-2-27B-gguf-dev" itself is not intended for regular dialogue or practical use. It is reported to be a development build released for the purpose of kernel development and operational verification toward upstream integration into llama.cpp.
How It Differs from Similar Models
Unlike the already released GGUF version for general use 三値量子化された27Bモデル「Bonsai 2 27B GGUF」公開, this model is said to be a test build until upstream support in llama.cpp is completed. It differs in that running it on standard llama.cpp generates nonsensical output without warnings, making operation on Prism ML’s fork of llama.cpp a prerequisite. It is also reported to differ in expected platforms and operating requirements from Qwen3.8-27Bベースの3値ウェイトモデル「Bonsai 2 27B」MLX 2bit版公開, which is optimized for Apple Silicon environments.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 27.8B parameters (taken from the base model Qwen/Qwen3.8-27B)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 12GB (RTX 4070 / 3060 12GB, etc.) | Q2_0 | 7.1GB | 8.5GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Recent Models in the Same Size Class
Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| Edge0/Edge0-35B-A3B-preview | 34.7B | 80GB | apache-2.0 | Edge0-35B-A3B-Preview: Sparse MoE for Phone-Class Memory (2026-09-11) |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | 12GB | apache-2.0 | Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations (2026-09-11) |
| nex-agi/Nex-N2.5-mini | 35.1B | 16GB | apache-2.0 | Nex-AGI Releases Open-Weight Long-Task Model Nex-N2.5-mini (2026-09-09) |
How to Get It
- Distribution format: GGUF
This model is a test build and requires the llama.cpp fork provided by Prism ML to run.
For regular use cases, it is recommended to use Ternary-Bonsai-2-27B-gguf or Ternary-Bonsai-2-27B-mlx-2bit for Apple Silicon environments. Setup instructions for each backend are provided in Bonsai-demo.
Related Articles
- Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model
- Signal-3.8-27B-GGUF: Faster and More Token-Efficient
- Salience-27B-R6 GGUF Released by bartowski
Sources
This article contains unverified information. We will append an update note once it is confirmed by a primary source.

