SQBench: leaderboard
Metric: Weighted Strict Pass@1 (%; 0.2 times L1 plus 0.6 times L2 plus 0.2 times L3 strict pass rates, the prespecified primary aggregate on SQBench v1.0, 220 production-oriented delivery tasks in which an agent processes input assets, uses tools and produces a specified deliverable; a task passes only with functional Completion 1 and no triggered 10-dimension risk penalty; one run per configuration and task; higher is better). Source: arxiv.org. Saturation forecast: Around September 2027. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K3 | 60.5 |
| 2 | Claude Opus 4.8 (Max) | 57.6 |
| 3 | GLM-5.2 (Max) | 52.7 |
| 4 | Qwen 3.7 Max | 51.1 |
| 5 | Kimi K2.7 Code | 50.4 |
| 6 | DeepSeek V4 Pro (Max) | 49.5 |
| 7 | MiMo-V2.5 | 48.5 |
| 8 | DeepSeek V4 Pro (High) | 48.3 |
| 9 | Doubao-Seed-2.1-Pro (High) | 48.1 |
| 10 | Kimi K2.6 | 47.3 |
| 11 | DeepSeek V4 Flash (High) | 47.1 |
| 12 | GLM-5.1 | 46.7 |
| 13 | MiniMax-M3 | 46.7 |
| 14 | Qwen 3.6 Plus | 45.4 |
| 15 | Qwen 3.7 Plus | 45.3 |
Interactive version: theaggregate.ai/benchmark?slug=sqbench · How It Works · Data refreshed daily, snapshot 2026-09-29.