

Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
Laguna S 2.1 is a 118B total parameter Mixture-of-Experts (MoE) model with 8B activated parameters per token and supports a context window of up to 1M tokens in thinking and no-thinking modes. It went from the start of training to launch in under nine weeks, and on long-horizon coding benchmarks it holds its own against models many times its size. For every benchmark score we publish today, we are releasing full trajectories for every trial in the final evaluation set at trajectories.poolside.ai .
Laguna S 2.1 118B-A8B Tencent Hy3 295B-A21B Inkling 975B-A41B Nemotron 3 Ultra 550B-A55B DeepSeek-V4-Pro-Max 1.6T-A49B Kimi K3 2.8T-A50B Qwen 3.7 Max — Muse Spark 1.1 — Claude Fable 5 — Terminal-Bench 2.1 Terminal-Bench 2.1 Resolved tasks on Terminal-Bench 2.1. 0.0 0.0 0.0 0.0 0 0.0 0.0 0 0 SWE-Bench Multilingual SWE-Bench Multilingual Resolved tasks on SWE-Bench Multilingual. 0.0 0.0 0.0 0.0 0.0 SWE-Bench Pro (Public Dataset) SWE-Bench Pro (Public Dataset) Resolved tasks on SWE-Bench Pro (Public Dataset). 0.0 0.0 0.0 0.0 0.0 0.0 0.0 DeepSWE DeepSWE Resolved tasks on DeepSWE. 0.0 0 0 0.0 0 SWE Atlas (Codebase QnA) SWE Atlas (Codebase QnA) Resolved tasks on SWE Atlas (Codebase QnA). 0.0 0.0 0.0 Toolathlon Verified Toolathlon Verified Resolved tasks on Toolathlon Verified. 0.0 0.0 0.0 0.0 0.0 Benchmarks as of 21 July 2026. pass@1 averaged over 4 attempts per task, except DeepSWE, SWE Atlas (Codebase QnA) and Toolathlon Verified that had 3 attempts per task. For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis), except SWE Atlas (Codebase QnA) where we do not use third-party leaderboard figures. Laguna S 2.1 (118B-A8B) Tencent Hy3 (295B-A21B) Inkling (975B-A41B) Nemotron 3 Ultra (550B-A55B) DeepSeek-V4-Pro Max (1.6T-A49B) Kimi K3 (2.8T-A50B) Qwen 3.7 Max (—) Muse Spark 1.1 (—) Claude Fable 5 (—) Terminal-Bench 2.1 70.2 71.7 63.8 56.4 64.0 88.3 74.5 80 88.0 SWE-Bench Multilingual 78.5 75.8 - 67.7 76.2 - 78.3 - - SWE-Bench Pro (Public Dataset) 59.4 57.9 54.3 - 55.4 - 60.6 61.5 80.3 DeepSWE 40.4 - - - 9.0 69.0 - 53.3 70.0 SWE Atlas (Codebase QnA) 46.2 - - - 27.2 - - 42.2 - Toolathlon Verified 49.7 - 45.5 34.3 55.9 - - 75.6 - Punching above its weight class Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class by a wide margin .
S 2.1 scores 70.2% on Terminal-Bench 2.1 in our agent harness with thinking enabled. Its compact size makes it uniquely suitable for complex work on local machines.
Terminal-Bench 2.1 evaluates a wide, high-quality set of long-horizon tasks where an agent model is connected to its environment through a terminal. Laguna S 2.1 is a standout model in its size category on this benchmark.
Laguna S 2.1 Laguna S 2.1: 70.2 at 118B Laguna XS 2.1 Laguna XS 2.1: 33.4 at 33B Kimi K3 Kimi K3: 88.3 at 2800B DeepSeek-V4-Pro-Max DeepSeek-V4-Pro-Max: 64 at 1600B Inkling Inkling: 63.8 at 975B Nemotron 3 Ultra Nemotron 3 Ultra: 56.4 at 550B MiniMax M3 MiniMax M3: 66 at 428B Hy3 Hy3: 71.7 at 295B DeepSeek-V4-Flash-Max DeepSeek-V4-Flash-Max: 61.8 at 284B Inkling-Small Inkling-Small: 52.7 at 276B Nemotron 3 Super Nemotron 3 Super: 38.6 at 120B Mistral Small 4 Mistral Small 4: 21.4 at 119B Qwen3.6-35B-A3B Qwen3.6-35B-A3B: 44.9 at 35B Qwen3.6-27B Qwen3.6-27B: 51.3 at 27B Total parameters (B, log scale) Score (Pass@1) Total parameters, log scale. Models with undisclosed total parameter counts are omitted. A closer look at DeepSWE The benchmarks above are all meaningful, and we're glad to be close to the frontier on them. But part of that closeness is a property of maturing benchmarks: as the frontier advances, top scores cluster in the 70-90% range and models that behave very differently end up no more than a few points apart. Datacurve’s DeepSWE still has significant headroom. Its tasks are longer-horizon and hard to partially solve, and the scores actually spread: frontier models range from 54% to 73% on the v1.1 variant, with some 1T+ parameter open models scoring below 10%.
On DeepSWE v1.1, Laguna S 2.1 scores 40.4 in thinking mode in pool harness.
It is worth noting that Laguna S 2.1 scored 40.4% in our agent harness, pool, not mini-swe-agent which DeepSWE’s leaderboard uses. For other models we report maximal over reported scores which for most models are the official leaderboard results reported by Datacurve. While this makes scores less comparable, we don’t believe it puts us in a particularly advantageous position as it’s been reported that many of the models score the same or better in mini-swe-agent compared to their native harnesses. Every trajectory in the final evaluation run is available here .
Evaluation of agent models is notoriously difficult due to prevalence of reward hacking. We have previously written about reward hacking in leading benchmarks and our evaluations system and rigor as part of the technical report on our Laguna M.1 and XS.2 models. Recent work has focused on adversarial judging to increase reward hacking detection.
With this release, we are making all trajectories from our final evaluations of the published Laguna S 2.1 checkpoint available to view and download at trajectories.poolside.ai .
Benchmark scores give a quantitative view into the model behavior, but to get a better intuitive understanding of how the model works it’s useful to look into runs on real world tasks. We share three such tasks with unedited trajectories and commentary.
One of our favorite things about Laguna S 2.1 is its resourcefulness : It will find clever ways to get to the goal even if the direct path is not available. We saw a great demonstration of this when we asked it to build a browser engine from scratch; knowing it would be a challenge for Laguna to verify its work given its lack of vision capabilities. In one 50-minute session of 181 steps, with no human intervention, Laguna S 2.1 built a working HTML/CSS rendering engine from an empty folder, then proved it renders like a real browser by measuring itself against one. Throughout its work, the model found increasingly complex ways to validate its work despite its limitations, leading to running headless Chromium to read canvases back and comparing screenshots numerically. See the full trajectory here .
// the verbatim prompt · reproduce it yourself your job is it to build a simple browser engine (just html/css) in javascript to demonstrate the capabilities of poolsides new "Laguna S" model. the goal is to take render html snippets in a canvas like a real browser. to demonstrate it the engine, build a self-contained single page app that showcases a gallery of multiple html snippets and renders them side by side (canvas with our render engine + iframe letting the hosting browser render it for real for comparison). support for most common layout and styling elements Over the session the model built the full pipeline, parser → cascade → layout → renderer, in vanilla JavaScript: an HTML tokenizer and DOM tree, a CSS parser with selector specificity, a cascade engine with inheritance, box-model layout, and a canvas-2D renderer, wrapped in an app that shows nine snippets on its own canvas beside the same markup in an iframe, so the hosting browser sits right there as the reference.
Hacker News
news.ycombinator.com