BenchGecko Score: leaderboard
Metric: Composite Score. Source: benchgecko.ai. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 Pro | 99.9 |
| 2 | Claude Mythos Preview | 97.8 |
| 3 | Qwen 3.5 397B A17B | 96.3 |
| 4 | DeepSeek V3.2 Speciale | 95.2 |
| 5 | Claude Fable 5 | 93.2 |
| 6 | GLM-5.2 | 92.7 |
| 7 | Step 3.5 Flash | 89.5 |
| 8 | GPT-5 Chat | 89 |
| 9 | GPT-5.4 Pro (xHigh) | 87.9 |
| 10 | DeepSeek V4 Pro | 86.5 |
| 11 | DeepSeek V2 | 86.1 |
| 12 | GPT-5.1 Codex Max | 85.6 |
| 13 | DeepSeek R1 Distill Qwen 14B | 85.4 |
| 14 | Claude Instant | 84.6 |
| 15 | Claude Opus 4.8 | 82.8 |
Interactive version: theaggregate.ai/benchmark?slug=benchgecko-score · How It Works · Data refreshed daily, snapshot 2026-09-05.