Skip to main content
Laguna XS.2 poolside laptop achieving 30.1% Terminal-Bench 2.0 score, outperforming Haiku 4.5 in performance comparison

Editorial illustration for Poolside's free Laguna XS.2 scores 30.1% on Terminal‑Bench 2.0, edging Haiku 4.5

Poolside's free Laguna XS.2 scores 30.1% on...

Updated: 3 min read

Poolside dropped Laguna XS.2 on Tuesday—a free model packing just 3 billion active parameters. Its score? A 30.1% on Terminal-Bench 2.0.

That’s a hair above Claude Haiku 4.5’s 29.8%. More strikingly, on SWE-bench Pro, this compact release outperformed not only Haiku but also the much larger, 31-billion-parameter Gemma 4 dense model. Specialized nano models, like GPT-5.4 Nano with its leading 46.3% terminal score, still hold the crown.

Despite having only 3B active parameters, XS.2 surpasses Claude Haiku 4.5 (39.5%) and the significantly larger Gemma 4 31B dense model (35.7%) on SWE-bench Pro. In terminal-based reasoning, XS.2’s 30.1% on Terminal-Bench 2.0 also edges out Haiku 4.5’s 29.8%, although it remains behind specialized "nano" models such as GPT-5.4 Nano, which reached 46.3% on the same benchmark. Collectively, these benchmarks suggest that Poolside’s focus on agentic RL and synthetic data curation has allowed its smaller models to "punch up" into weight classes typically reserved for far denser architectures.

Laguna XS.2’s performance, beating Haiku, signals a shift. Poolside’s recipe—agentic reinforcement learning and curated synthetic data—is proving you don’t need sheer parameter volume. The model still trails far behind GPT-5.4 Nano’s 46.3% on terminals. Yet for a free, locally run option with only 3 billion parameters, outpacing larger commercial peers redefines what’s possible at this scale.

Common Questions Answered

How does Poolside's Laguna XS.2 compare to Claude Haiku 4.5 on Terminal-Bench 2.0?

Laguna XS.2 scored 30.1% on Terminal-Bench 2.0, slightly edging out Claude Haiku 4.5's 29.8% score. This is particularly impressive given that Laguna XS.2 is a free model with only 3 billion active parameters, demonstrating competitive performance despite its smaller size and no licensing costs.

What makes Laguna XS.2 notable on SWE-bench Pro compared to larger models?

On SWE-bench Pro, Laguna XS.2 outperformed both Claude Haiku 4.5 and Gemma 4, which has 31 billion parameters—over 10 times larger than Laguna XS.2. This achievement demonstrates that parameter count alone doesn't determine performance, as Poolside's approach using agentic reinforcement learning and curated synthetic data proved more effective.

What techniques does Poolside use to achieve strong performance with only 3 billion parameters?

Poolside employs agentic reinforcement learning and curated synthetic data as its core recipe for training Laguna XS.2. These techniques enable the model to punch above its weight class, allowing a compact 3 billion parameter model to outperform much larger commercial peers and redefine what's possible at this scale.

Which model currently holds the leading score on Terminal-Bench 2.0?

GPT-5.4 Nano holds the crown with a leading 46.3% score on Terminal-Bench 2.0. While Laguna XS.2's 30.1% trails significantly behind this specialized nano model, it still represents a strong performance for a free, locally-runnable option with minimal parameters.

LIVE21:05Delhi High Court Rejects News Agency's Copyright Injunction Against OpenAI