STT-Arena - Easy: leaderboard

Metric: Pass@1 (%) on the 71 easy tasks (a single isolated conflict needing one corrective action); a task passes when every rule-based check function holds in the final state; the model acts in an executable stateful environment whose state changes under scheduled spatio-temporal triggers, with a passive Qwen3.5-397B user simulator and at most 50 turns; mean of three runs at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.641.78
2GPT-5.439.91
3Gemini 3.1 Pro (Preview)37.56
4DeepSeek V3.236.62
5Claude Sonnet 4.635.21
6Qwen 3.6 Plus35.21
7GLM-5.135.21
8DeepSeek V4 Pro32.86
9GLM-530.05
10Gemini 2.5 Flash28.17
11Qwen 3.5 9B27.23
12MiniMax-M2.526.76
13MiniMax-M2.726.29
14Kimi K2.525.35
15Llama 3.3 70B22.07

Interactive version: theaggregate.ai/benchmark?slug=stt-arena-easy · How It Works · Data refreshed daily, snapshot 2026-10-07.