STT-Arena - Impossible: leaderboard

Metric: Pass@1 (%) on the 30 impossible tasks (no valid completion path exists); a task passes when a Qwen3.5-397B judge finds that the model both recognized the infeasibility and told the user; the model acts in an executable stateful environment whose state changes under scheduled spatio-temporal triggers, with a passive Qwen3.5-397B user simulator and at most 50 turns; mean of three runs at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.

Top models

#ModelScore
1DeepSeek V3.263.33
2Claude Sonnet 4.658.89
3Qwen 3.6 Plus54.44
4Gemini 3.1 Pro (Preview)52.22
5GPT-5.452.22
6Claude Opus 4.652.22
7DeepSeek V4 Pro45.56
8GLM-543.33
9MiniMax-M2.542.22
10GLM-5.141.11
11Kimi K2.537.78
12MiniMax-M2.736.67
13Qwen 3.5 9B32.22
14Gemini 2.5 Flash31.11
15Llama 3.3 70B30

Interactive version: theaggregate.ai/benchmark?slug=stt-arena-impossible · How It Works · Data refreshed daily, snapshot 2026-10-07.