Skip to content
← All news
4 min read

Poolside's new open model beats models six times its size at agentic coding

Laguna S 2.1 is a 118B parameter model, 8B active, that scored 78.5 percent on SWE-Bench Multilingual and runs on a single desktop GPU box.

Poolside's 118B open model beats bigger rivals at agentic coding, runs on one GPU box.

Poolside released Laguna S 2.1 on July 21, an open weight, 118 billion parameter model with only 8 billion active per token, built specifically for agentic coding and long horizon tasks.

The benchmark numbers

Laguna S 2.1 scores 70.2 percent on Terminal-Bench 2.1, 78.5 percent on SWE-Bench Multilingual, topping the published table outright, and 40.4 percent on DeepSWE v1.1 against DeepSeek-V4-Pro-Max's 9.0 percent, despite running with roughly one sixth the active parameters. It is the third release in Poolside's self imposed "three models in three months" cadence, trained on 4,096 Nvidia H200 GPUs over less than nine weeks.

What actually changed under the hood

Same pretraining data as the prior Laguna XS 2.1 release, but Poolside ran reinforcement learning in FP8 precision for the first time, added a new sandboxing service, and trained across multiple agent harnesses instead of one, specifically to avoid overfitting to a single scaffold. Poolside co-head of applied research Pengming Wang put the goal plainly: "not necessarily add more intelligence, but improve the behaviors that lead to a more capable model: more verification, less taking things for granted, not declaring victory early, and being more persistent."

Why a build studio cares

A coding-focused open model that fits on a single desktop GPU box and beats larger rivals on real benchmarks is directly relevant to any build weighing self hosted agentic coding tools against API-only options. Worth testing against an actual repo before trusting the benchmark table, same caveat as any model release, but the resource footprint claim, one GPU box, is the part worth verifying first since it is what makes this practically different from the alternatives.

Next step: read Poolside's own writeup. If you're evaluating open weight coding models for a build, write to us at hello@gattyworks.com.

AIDev ToolsOpen SourcePoolsideAILagunaS21AgenticCodingOpenWeightAISWEBenchAICodeGenLLMDevToolsMachineLearningOpenSource

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.