SQBench - L1 Atomic Capabilities: leaderboard

Metric: Strict Pass@1 (%; the 100 L1 atomic-capability tasks on SQBench v1.0, 220 production-oriented delivery tasks in which an agent processes input assets, uses tools and produces a specified deliverable; a task passes only with functional Completion 1 and no triggered 10-dimension risk penalty; one run per configuration and task; higher is better). Source: arxiv.org. Saturation forecast: Around July 2027. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (Max)63
2Kimi K361
3Qwen 3.7 Max52
4GLM-5.2 (Max)50
5Gemini 3.5 Flash (High)49
6DeepSeek V4 Pro (Max)46
7Qwen 3.7 Plus45
8Hy343
9Qwen 3.6 Plus42
10GLM-5.142
11Kimi K2.7 Code42
12MiMo-V2.541
13Kimi K2.640
14DeepSeek V4 Pro (High)40
15Doubao-Seed-2.1-Pro (High)39

Interactive version: theaggregate.ai/benchmark?slug=sqbench-l1-atomic-capabilities · How It Works · Data refreshed daily, snapshot 2026-09-29.