SuperCLUE General (July 2026) - Agentic Task Planning: leaderboard
Metric: Score. Source: www.superclueai.com. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.3 (Max) | 91.15 |
| 2 | Qwen 3.8 Max (Max) | 90.94 |
| 3 | Qwen 3.8 Flash (Max) | 88.61 |
| 4 | GLM-5.3 Flash (Max) | 88.43 |
| 5 | DeepSeek V4 Pro (0813) (Max) | 85.48 |
| 6 | Kimi K3 (Max) | 84.35 |
| 7 | DeepSeek V4 Flash (Max) | 80.72 |
| 8 | DeepSeek V4.1 Flash (Max) | 79.84 |
| 9 | DeepSeek V4 Pro (Max) | 79 |
| 10 | MiniMax-M3 | 77.68 |
| 11 | LongCat 2.0 | 71.32 |
| 12 | Hy3 (High) | 66.82 |
| 13 | Step 3.7 Flash (High) | 66.54 |
| 14 | GLM-5.2 (Max) | 65.27 |
| 15 | MiMo-V2.5-Pro | 57.25 |
Interactive version: theaggregate.ai/benchmark?slug=superclue-general-july-2026-agentic-task-planning · How It Works · Data refreshed daily, snapshot 2026-09-19.