Editorial illustration for Poolside AI launches Laguna XS.2 and M.1, hitting 72.5% on SWE-bench Verified
Poolside AI launches Laguna XS.2 and M.1, hitting 72.5%...
Benchmarks are a game of inches, until they’re not. Poolside AI just took a yard. Its new Laguna XS.2 and M.1 models scored 68.2% and 72.5% on SWE-bench Verified.
The bigger number is flashy. The smaller one is the trapdoor. XS.2 is a 33-billion-parameter model that activates only 3 billion per token.
It’s small enough to run locally on a 36GB Mac. This is no lab experiment. It’s a tool you can install.
The M.1 model, presumably larger, posted wider scores: 67.3% on SWE-bench Multilingual, 46.9% on the Pro version, 40.7% on Terminal-Bench 2.0. XS.2 hits 62.4%, 44.5%, and 30.1% on those same tests. The architecture is built for this efficiency.
It uses sigmoid gating, per-layer rotary scales, and mixes sliding window with global attention in a 3:1 ratio across 40 layers. Poolside is also releasing an XS.2-base variant for anyone who wants to fine-tune it themselves.
Poolside AI released the first two models in its Laguna family: Laguna M.1 and Laguna XS.2 .
A 72.5% score is a technical victory. It’s a line in a press release. The ability to put a model that scores 68.2% on a developer’s laptop is something else entirely.
It’s a shift in control. The meticulous architecture, the gating, the attention mix, they aren’t academic details. They are the engineering that makes local, capable coding agents physically possible.
M.1 shows the ceiling. XS.2 shows the door is open. The promise of open-weight, agentic AI is slowly hardening into a file you can run.
Common Questions Answered
What are the SWE-bench Verified scores achieved by Poolside AI's Laguna models?
Poolside AI's Laguna XS.2 model scored 68.2% on SWE-bench Verified, while the larger Laguna M.1 model achieved 72.5% on the same benchmark. These scores represent a significant advancement in software engineering task performance for the Laguna model line.
How does the Laguna XS.2 model's parameter efficiency work with its 33-billion-parameter design?
The Laguna XS.2 is a 33-billion-parameter model that uses a gating mechanism to activate only 3 billion parameters per token, making it efficient enough to run locally on a 36GB Mac. This architectural approach allows for capable performance while maintaining a small computational footprint suitable for local deployment.
What is the significance of running Laguna XS.2 on a developer's laptop?
The ability to run Laguna XS.2 locally on a 36GB Mac represents a shift in control, enabling developers to deploy capable coding agents on their own hardware without relying on cloud infrastructure. This accessibility makes open-weight, agentic AI physically possible for individual developers and organizations.
How do the Laguna XS.2 and M.1 models demonstrate different aspects of Poolside AI's capabilities?
The Laguna M.1 with its 72.5% SWE-bench score shows the performance ceiling of what Poolside AI can achieve, while the XS.2 demonstrates that capable models can be made accessible for local deployment. Together, they illustrate both the technical ceiling and the practical accessibility of open-weight coding AI models.