SQBench - L3 Business Scenarios: leaderboard

Metric: Strict Pass@1 (%; the 60 L3 domain business-scenario tasks on SQBench v1.0, 220 production-oriented delivery tasks in which an agent processes input assets, uses tools and produces a specified deliverable; a task passes only with functional Completion 1 and no triggered 10-dimension risk penalty; one run per configuration and task; higher is better). Source: arxiv.org. Saturation forecast: Around 2030. 27 models tracked.

Top models

#ModelScore
1GLM-5.2 (Max)28.3
2Kimi K326.7
3Claude Opus 4.8 (Max)25
4DeepSeek V4 Flash (High)23.3
5Kimi K2.621.7
6GLM-5.121.7
7Qwen 3.7 Plus21.7
8MiMo-V2.521.7
9DeepSeek V4 Pro (Max)21.7
10DeepSeek V4 Pro (High)21.7
11MiMo-V2.5-Pro20
12Kimi K2.7 Code20
13DeepSeek V4 Flash (Max)20
14MiniMax-M2.718.3
15MiniMax-M318.3

Interactive version: theaggregate.ai/benchmark?slug=sqbench-l3-business-scenarios · How It Works · Data refreshed daily, snapshot 2026-09-29.