Laguna M.1: benchmark results
Poolside's open 225B MoE (23B active) agentic coding model trained with code-execution RL, scoring 72.5% on SWE-bench Verified (April 2026). Provider: Poolside. Released 2026-04-28. Access: Open.
Unified ELO 1563 ± 1, rank #392 of 1392 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Continuation) | 743 | Elo | 66.7 |
| AI Chess Leaderboard (Reasoning) | 769 | Elo | 62.5 |
| Wolfram LLM Benchmarking Project | 46 | Correct Functionality (%) | 59.5 |
| Vals AI Terminal-Bench 2.0 | 31.46 | Accuracy (%) | 36.4 |
| Vals AI CorpFin v2 | 58.16 | Accuracy (%) | 35.3 |
| Vals AI LiveCodeBench | 68.12 | Accuracy (%) | 34.5 |
| SvelteBench | 80 | Average pass@1 (%) | 27.4 |
| Chatbot Arena (Code) | 1347.49 | Elo | 24.4 |
| WebDev Arena (Frontend) | 1342 | Arena Score | 24.4 |
| WebDev Arena (Reference-Based Design) | 1346 | Arena Score | 23 |
| WebDev Arena | 1347.31 | Arena Score | 22.2 |
| Vals AI LegalBench | 75.14 | Accuracy (%) | 22 |
Interactive version: theaggregate.ai/model?slug=laguna-m-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.