STT-Arena - Medium: leaderboard

Metric: Pass@1 (%) on the 67 medium tasks; a task passes when every rule-based check function holds in the final state; the model acts in an executable stateful environment whose state changes under scheduled spatio-temporal triggers, with a passive Qwen3.5-397B user simulator and at most 50 turns; mean of three runs at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 23 models tracked.

Top models

#ModelScore
1Claude Opus 4.631.34
2Gemini 3.1 Pro (Preview)30.35
3Claude Sonnet 4.629.85
4GPT-5.429.85
5DeepSeek V3.227.86
6Qwen 3.6 Plus24.38
7GLM-5.123.38
8GLM-521.39
9DeepSeek V4 Pro20.4
10MiniMax-M2.715.92
11MiniMax-M2.515.92
12Gemini 2.5 Flash13.43
13Kimi K2.513.43
14Qwen 3 8B13.43
15GPT-5.4 Mini13.43

Interactive version: theaggregate.ai/benchmark?slug=stt-arena-medium · How It Works · Data refreshed daily, snapshot 2026-10-07.