SQBench: leaderboard

Metric: Weighted Strict Pass@1 (%; 0.2 times L1 plus 0.6 times L2 plus 0.2 times L3 strict pass rates, the prespecified primary aggregate on SQBench v1.0, 220 production-oriented delivery tasks in which an agent processes input assets, uses tools and produces a specified deliverable; a task passes only with functional Completion 1 and no triggered 10-dimension risk penalty; one run per configuration and task; higher is better). Source: arxiv.org. Saturation forecast: Around September 2027. 27 models tracked.

Top models

#ModelScore
1Kimi K360.5
2Claude Opus 4.8 (Max)57.6
3GLM-5.2 (Max)52.7
4Qwen 3.7 Max51.1
5Kimi K2.7 Code50.4
6DeepSeek V4 Pro (Max)49.5
7MiMo-V2.548.5
8DeepSeek V4 Pro (High)48.3
9Doubao-Seed-2.1-Pro (High)48.1
10Kimi K2.647.3
11DeepSeek V4 Flash (High)47.1
12GLM-5.146.7
13MiniMax-M346.7
14Qwen 3.6 Plus45.4
15Qwen 3.7 Plus45.3

Interactive version: theaggregate.ai/benchmark?slug=sqbench · How It Works · Data refreshed daily, snapshot 2026-09-29.