Skip to main content
PrismML Ternary Bonsai 2 model, 5.9GB, matching Qwen3.8 performance, displayed on a screen with data visualizations.

Editorial illustration for PrismML's Ternary Bonsai 2 Model: 5.9 GB, Matches Qwen3.8 Performance

Ternary Bonsai 2: 5.9GB Model Matches Qwen Performance

4 min read

PrismML has shipped Ternary Bonsai 2 27B, a compressed rework of Qwen3.8 27B that fits in 5.93 GB instead of the original's 53.80 GB in FP16. The company says the model holds onto 98.2% of the parent's average score across 20 benchmarks, a jump from the roughly 95% retained by the first Bonsai 27B released two months ago. The model handles both text and images, supports a 262K-token context window, and PrismML has shown it running Cline coding agents and computer-use tasks on a single RTX 5090.

The weights are Apache 2.0 licensed and already runnable, PrismML says, on a 16 GB laptop or a 24 GB GPU, provided you're using its own llama.cpp fork or MLX runtime. The architecture itself is untouched from Qwen3.8 27B: 27.36B parameters split across a 24.35B backbone, 2.54B in embeddings and the LM head, and a 0.47B vision tower that loads separately as a 0.63 GB GGUF file only when an image shows up. The backbone runs hybrid attention, mixing linear and full-attention layers. The real story is how PrismML squeezed that footprint down in the first place, and it starts with how the ternary format itself represents each weight.

PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB, against 53.80 GB in FP16. PrismML reports that it keeps 98.2% of the parent model’s average across 20 benchmarks.

Why this matters

A 27B model that fits in 5.9 GB and still holds 98.2% of its parent's benchmark average changes what "local deployment" means in practice. Two months ago the first Bonsai gave up 5% of performance for the same trick; PrismML has cut that loss by more than half while keeping a 262K context window and image input intact. That's the number worth watching, not the raw file size. If ternary quantization keeps closing the gap this fast, the argument for renting cloud GPUs to run mid-size models gets weaker every quarter.

The RTX 5090 numbers (142.5 tokens/sec at 0.582 mWh per token) matter more to founders than researchers. That's power efficiency good enough to put agentic coding tools like Cline on a desktop tower instead of a metered API. The 4090 and even a 72W L4 still clearing 32 tokens/sec on the same weights suggests this scales down to edge boxes, not just enthusiast rigs.

Apache 2.0 licensing means nothing here is locked behind a partnership. The real test is whether PrismML's kernels generalize past Qwen3.8, or whether this is one lucky architecture.

Common Questions Answered

How much smaller is Ternary Bonsai 2 27B compared to the original Qwen3.8 27B model?

Ternary Bonsai 2 27B is compressed to just 5.93 GB, compared to the original Qwen3.8 27B's 53.80 GB in FP16 format. This represents approximately a 9x reduction in file size while maintaining the model's core functionality and performance characteristics.

What percentage of performance does Ternary Bonsai 2 retain compared to its parent model?

Ternary Bonsai 2 retains 98.2% of the parent Qwen3.8 27B's average score across 20 benchmarks. This is a significant improvement over the first Bonsai 27B released two months earlier, which retained only approximately 95% of the parent model's performance.

What capabilities does Ternary Bonsai 2 support despite its compressed size?

Ternary Bonsai 2 supports both text and image inputs, maintains a 262K-token context window, and can run complex tasks like Cline coding agents and computer-use applications on a single RTX 5090 GPU. These capabilities demonstrate that the model's compression doesn't sacrifice multimodal functionality or advanced reasoning tasks.

How much has the performance loss improved between the first and second versions of Bonsai?

The performance loss has been cut by more than half between Bonsai 27B and Ternary Bonsai 2 27B. While the first version lost approximately 5% of performance, the second version loses only 1.8%, showing rapid improvement in PrismML's ternary quantization technique.

LIVE00:06Anthropic Says Claude Leads Research, But Its AI Judge Could Repeat Errors