Skip to main content
Poolside's Laguna S 2.1 coding model, a leading open-weight AI, displayed on a screen with performance metrics.

Editorial illustration for Poolside's Laguna S 2.1 Coding Model Leads Open-Weight Pack on SWE-Bench

Laguna S 2.1 Tops Open-Weight Coding Models

4 min read

Poolside shipped Laguna S 2.1 this week, a 118-billion-parameter open-weight model built for agentic coding that runs on a single NVIDIA DGX Spark. That's the headline number worth sitting with, because the model only activates about 8 billion parameters per token, a Mixture-of-Experts setup that lets it punch well above its footprint on long-horizon coding tasks. Pre-training started 22 May 2026 on 4,096 H200 GPUs, and Poolside got from that starting line to launch in under nine weeks, with reinforcement learning running in FP8 precision for the first time on any Poolside model.

The weights are live on Hugging Face under an OpenMDW-1.1 license, with BF16, FP8, INT4, and NVFP4 versions plus GGUF and MLX conversions for anyone who wants to run it locally. It supports context windows up to 1 million tokens in both thinking and no-thinking modes, and it's a scale-up of the Laguna XS line, trained on the same pre-training data as XS 2.1. On the benchmarks that matter for agentic coding, Poolside is putting this mid-size model in the same conversation as DeepSeek-V4-Pro-Max and Nemotron 3 Ultra.

Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled. That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems. On SWE-Bench Multilingual it scores 78.5%, topping the published table outright.

Why this matters

An 8B-active-parameter MoE beating models several times its size on SWE-Bench Multilingual and Terminal-Bench 2.1 is the kind of result that should make anyone budgeting GPU hours pay attention. Poolside is betting that a well-trained mixture of experts, plus a 1M-token context window, closes the gap that used to require sheer scale. That's a meaningful shift for teams building coding agents who don't have Anthropic or OpenAI's compute budget.

Running on a single DGX Spark and shipping under OpenMDW-1.1 also matters for founders who want to fine-tune or audit a model without renegotiating licensing terms every quarter. Still, one benchmark cycle doesn't settle the argument. Poolside compiled its own leaderboard, and "first among open, disclosed-size models" is a narrower claim than "best coding model." We'd want to see Laguna S 2.1 tested on private repos and messier, real-world codebases before crowning it.

For researchers, though, the activated-parameter efficiency here is worth studying regardless of where it lands on any single chart. Smaller active-parameter counts with strong long-horizon performance point toward cheaper agentic coding, and that's the trend to watch.

Common Questions Answered

What is the parameter efficiency of Poolside's Laguna S 2.1 model?

Laguna S 2.1 is a 118-billion-parameter open-weight model that only activates about 8 billion parameters per token through a Mixture-of-Experts setup. This efficient architecture allows the model to perform significantly better on long-horizon coding tasks while maintaining a smaller computational footprint than its total parameter count suggests.

How does Laguna S 2.1 perform on SWE-Bench Multilingual compared to other models?

Laguna S 2.1 scores 78.5% on SWE-Bench Multilingual, topping the published table outright among disclosed models. This performance is particularly impressive given that the model only activates 8 billion parameters per token, demonstrating that efficient architecture can compete with much larger systems.

What hardware requirements does Laguna S 2.1 need to run?

Laguna S 2.1 runs on a single NVIDIA DGX Spark, making it accessible for teams without massive GPU infrastructure. This accessibility is significant for teams building coding agents who lack the compute budgets of larger AI companies like Anthropic or OpenAI.

How quickly was Laguna S 2.1 developed from pre-training to launch?

Poolside completed the development of Laguna S 2.1 in under nine weeks, starting pre-training on May 22, 2026 using 4,096 H200 GPUs. This rapid development timeline demonstrates efficient training and optimization practices for the agentic coding model.

What is the context window size for Laguna S 2.1?

Laguna S 2.1 features a 1 million-token context window, which combined with its Mixture-of-Experts architecture, helps close the performance gap that previously required significantly larger model scale. This extended context capability is particularly valuable for handling complex, long-horizon coding tasks.

LIVE04:12Meta Tests 'StoryKit' AI App for Children's Bedtime Stories