Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

At a Glance

Item Value
Repository prism-ml/Ternary-Bonsai-2-27B-gguf-dev
Published 2026-09-17
License apache-2.0
Formats GGUF
Source type Unverified (not confirmed by a primary source)

Values determined by this site’s code at collection time. Dates are JST.

Overview

It is reported that Prism ML has released a development and testing GGUF build, prism-ml/Ternary-Bonsai-2-27B-gguf-dev, for “Bonsai 2 27B", a ternary weight model based on Qwen3.8-27B.

However, this news has not been officially confirmed and is currently considered unconfirmed information. It is reported that this build was prepared for the purpose of kernel development and testing for Q2_0 packing in llama.cpp, as well as verification for future upstream integration.

Specifications

  • Parameters: 27B (base model)
  • Architecture: Qwen3_5ForConditionalGeneration (base model)
  • Context length: 262,144 tokens (base model)

Performance

Regarding the performance of the base model “Qwen3.8-27B", the comparison table provided in the model card from the publisher is as follows. Note that the comparison targets have been narrowed down to four.

Text Performance

→ Scroll horizontally to see all columns

Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Opus4.6 Max
Coding
Agentic terminal coding Terminal Bench 2.1 (Terminus) 73.0 63.4 64.0 78.2
Agentic coding SWE-bench Pro 61.7 53.5 57.6 53.4
Repo-level code generation NL2Repo-Bench 42.3 36.2 41.1 47.6
Agentic coding DeepSWE 1.1 42.2 13.3 14.2
Software engineering QwenSWEBench 79.0 49.3 59.2 63.8
Agent
Long-horizon office work CoWorkBench 70.7 61.0 65.1 68.2
Professional job tasks JobBench 33.4 21.8 27.6
Frontier agentic tasks Agents’ Last Exam Pass@1 20.4 Score 42.9 Pass@1 10.6 Score 27.3 Pass@1 13.2 Score 33.6
General
Instruction following IFBench 79.5 69.1 79.1 62.5
Scientific reasoning GPQA Diamond 89.2 87.8 90.3 91.3
Multidisciplinary reasoning HLE 30.8 24.0 34.7 40.0
Competitive coding LiveCodeBench v6 90.3 83.9 89.6 88.8

VL Performance

→ Scroll horizontally to see all columns

Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Opus4.6 Max
Agentic Multimodal Intelligence
Computer use OSWorld-Verified 84.3 63.9 73.3 72.7
Browser use WebArena-Verified 64.8 48.8 55.3
Mobile use AndroidWorld 81.9 70.3 81.0 62.0
Application recreation RecreationBench 47.1 29.8 30.2
Multimodal tool use ClawEval-MM Pass@3 57.4 Average 56.9 Pass@3 42.6 Average 50.4 Pass@3 57.4 Average 60.1 Pass@3 52.5 Average 54.7
Multimodal software engineering SWE-MM 38.6 25.7 30.0 27.1
Visual web development Vision2Web 62.9 45.0 42.1
General Multimodal Intelligence
Visual math problem solving MathVision Without CI 90.0 With CI 94.6 Without CI 85.1 Without CI 90.3 Without CI 65.5
General visual reasoning BabyVision Without CI 65.7 With CI 85.6 Without CI 28.9 Without CI 64.7 With CI 70.4 Without CI 12.6
Scientific chart analysis CharXiv (RQ) Without CI 83.7 With CI 90.2 Without CI 78.4 Without CI 85.8 With CI 85.9 Without CI 66.0
Document intelligence OmniDocBench 1.5 91.1 89.4 91.4 86.6
Real-world perception RealWorldQA 85.9 84.1 86.9 73.9
Embodied intelligence ERQA 65.5 62.5 69.8 40.8

According to the measurement results published by the base model’s creators, Qwen3.8-27B is reported to show high performance particularly in coding and agent tasks. It is reported to have recorded 61.7 on “SWE-bench Pro" (which measures the ability to resolve real GitHub issues) and 73.0 on “Terminal-Bench" (which measures terminal operation tasks), demonstrating capabilities that surpass “Opus4.6 Max" and “Qwen3.7-Plus" in some items. On the other hand, it is reported to fall behind in some extremely difficult tasks, remaining at 30.8 on “HLE" (which handles expert-level extremely difficult questions) compared to the top-tier model “Opus4.6 Max" at 40.0. Additionally, it scored 89.2 on “GPQA Diamond" (which measures scientific difficult questions), demonstrating high scientific reasoning capabilities approaching upper-tier models.

Strengths and Use Cases

The base model “Qwen3.8-27B" is reported to excel in coding, mathematics, long-context reading, multilingual support, agent tasks, and multimodal visual understanding (Vision-Language) including images and videos.

However, the newly released “Ternary-Bonsai-2-27B-gguf-dev" itself is not intended for regular dialogue or practical use. It is reported to be a development build released for the purpose of kernel development and operational verification toward upstream integration into llama.cpp.

How It Differs from Similar Models

Unlike the already released GGUF version for general use 三値量子化された27Bモデル「Bonsai 2 27B GGUF」公開, this model is said to be a test build until upstream support in llama.cpp is completed. It differs in that running it on standard llama.cpp generates nonsensical output without warnings, making operation on Prism ML’s fork of llama.cpp a prerequisite. It is also reported to differ in expected platforms and operating requirements from Qwen3.8-27Bベースの3値ウェイトモデル「Bonsai 2 27B」MLX 2bit版公開, which is optimized for Apple Silicon environments.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.8B parameters (taken from the base model Qwen/Qwen3.8-27B)

Your VRAM Quantization File size Est. memory needed
12GB (RTX 4070 / 3060 12GB, etc.) Q2_0 7.1GB 8.5GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Recent Models in the Same Size Class

Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
Edge0/Edge0-35B-A3B-preview 34.7B 80GB apache-2.0 Edge0-35B-A3B-Preview: Sparse MoE for Phone-Class Memory (2026-09-11)
bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF 26.5B 12GB apache-2.0 Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations (2026-09-11)
nex-agi/Nex-N2.5-mini 35.1B 16GB apache-2.0 Nex-AGI Releases Open-Weight Long-Task Model Nex-N2.5-mini (2026-09-09)

How to Get It

  • Distribution format: GGUF

This model is a test build and requires the llama.cpp fork provided by Prism ML to run.

For regular use cases, it is recommended to use Ternary-Bonsai-2-27B-gguf or Ternary-Bonsai-2-27B-mlx-2bit for Apple Silicon environments. Setup instructions for each backend are provided in Bonsai-demo.

Related Articles

Sources

This article contains unverified information. We will append an update note once it is confirmed by a primary source.