SQBench - L2 Composite Skills: leaderboard

Metric: Strict Pass@1 (%; the 60 L2 composite-skill tasks on SQBench v1.0, 220 production-oriented delivery tasks in which an agent processes input assets, uses tools and produces a specified deliverable; a task passes only with functional Completion 1 and no triggered 10-dimension risk penalty; one run per configuration and task; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 27 models tracked.

Top models

#ModelScore
1Kimi K371.7
2Claude Opus 4.8 (Max)66.7
3Kimi K2.7 Code63.3
4Qwen 3.7 Max61.7
5GLM-5.2 (Max)61.7
6Doubao-Seed-2.1-Pro (High)61.7
7MiniMax-M360
8MiMo-V2.560
9DeepSeek V4 Pro (Max)60
10DeepSeek V4 Pro (High)60
11DeepSeek V4 Flash (High)60
12Kimi K2.658.3
13Qwen 3.6 Plus56.7
14GLM-5.156.7
15MiMo-V2.5-Pro56.7

Interactive version: theaggregate.ai/benchmark?slug=sqbench-l2-composite-skills · How It Works · Data refreshed daily, snapshot 2026-09-29.