Skip to main content
Saluki 27B: Qwen3.8 in 2 bits, outperforming original at tool calling, shown on a computer screen with code.

Editorial illustration for Saluki 27B: Qwen3.8 in 2 Bits Outperforms Original at Tool Calling

Saluki 27B: 2-Bit Qwen Model Beats Original at Tool Calling

Saluki 27B: Qwen3.8 in 2 Bits Outperforms Original at Tool Calling

• 4 min read

Conway Research's Underdog project just put out Saluki 27B, a 2-bit compressed version of Qwen3.8-27B, under an Apache 2.0 license. The file comes in at 7.89 GB. The original BF16 model runs 54 GB, so this is roughly a 7x size cut on a dense 27-billion-parameter model.

That kind of compression usually costs you accuracy across the board. Conway's team instead tuned the quantization to protect one specific skill: tool calling, the mechanism that lets a language model act as an agent rather than just a chat partner. Across nine benchmarks, Saluki 27B holds onto 96% of the original model's performance on average, and it designed the model to run on stock llama.cpp with full GPU offload, plus an optional vision add-on under a gigabyte.

The tradeoffs aren't hidden. Some tasks hold up fine, others take real hits, and the gap between the two tells you exactly where this kind of compression pays off and where it doesn't. The breakdown, benchmark by benchmark, is below.

Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent.

Why this matters

For developers building on-device agents, Saluki 27B is a reminder that the right compression target matters more than the compression ratio itself. Conway Research didn't just shrink Qwen3.8-27B, they chose what to sacrifice. Tool calling went up, AIME math went down 17.5 points. That's a deliberate trade, not a side effect, and it's the kind of tuning decision that only makes sense if you know exactly what your model will be asked to do in production.

The 7.89 GB footprint running on stock llama.cpp is the practical headline: a 27B-class agent that fits on a laptop GPU or a beefy phone, no custom runtime required. But the AIME drop should temper any assumption that this is a free upgrade. If your pipeline leans on multi-step reasoning or math, this checkpoint will cost you.

Apache 2.0 licensing means teams can inspect and adapt Underdog's approach rather than trust a black box. Worth watching: whether other labs start publishing per-skill retention numbers instead of one blended benchmark score, because that's the only way to judge whether a compressed model fits your actual use case.

Common Questions Answered

How much smaller is Saluki 27B compared to the original Qwen3.8-27B model?

Saluki 27B is approximately 7 times smaller than the original model, reducing the file size from 54 GB (BF16 format) down to just 7.89 GB through 2-bit quantization compression. This dramatic size reduction makes the model significantly more practical for deployment on resource-constrained devices while maintaining strong performance in specific tasks.

What specific capability did Conway Research prioritize when compressing Saluki 27B?

Conway Research specifically tuned the quantization compression to protect tool calling, which is the mechanism that enables language models to act as agents rather than simple chat models. By prioritizing this skill during compression, they ensured that Saluki 27B actually outperforms the original Qwen3.8-27B at tool calling despite the aggressive 2-bit quantization.

What trade-offs were made in Saluki 27B's performance compared to the original model?

While tool calling performance improved, Saluki 27B experienced a deliberate trade-off in other areas, with AIME math performance declining by 17.5 points compared to the original model. This represents a conscious engineering decision by Conway Research to sacrifice performance in less critical tasks in order to maintain excellence in tool calling for on-device agent applications.

Under what license is Saluki 27B released?

Saluki 27B is released under the Apache 2.0 license, making it freely available for both commercial and non-commercial use. This open licensing approach enables developers to integrate the compressed model into their projects without restrictive licensing constraints.

Why is Saluki 27B's compression approach significant for developers building on-device agents?

Saluki 27B demonstrates that the most important factor in model compression is choosing the right capabilities to prioritize, not simply achieving the highest compression ratio possible. For developers building on-device agents, this means they can use a 7.89 GB model that excels at tool calling rather than forcing a larger model onto limited hardware, making it practical to deploy intelligent agents directly on user devices.

LIVE10:52Saluki 27B: Qwen3.8 in 2 Bits Outperforms Original at Tool Calling