SciExam-ENSO - Predictivity: leaderboard
Metric: Predictivity score (0–1), final submitted model, six-hour budget. Source: github.com. Saturation forecast: Around 2029. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol | 0.31 |
| 2 | Kimi K3 | 0.31 |
| 3 | Claude Fable 5.1 | 0.27 |
| 4 | GPT-5.5 | 0.25 |
| 5 | DeepSeek V4 Pro | 0.25 |
| 6 | Claude Opus 5 | 0.25 |
| 7 | Gemini 3.8 Flash | 0.24 |
| 8 | GPT-6 Astra | 0.24 |
| 9 | DeepSeek V4.1 Flash | 0.18 |
| 10 | MiniMax-M3 | 0.14 |
| 11 | GLM-5.2 | 0.11 |
| 12 | Qwen 3 Max | 0 |
Interactive version: theaggregate.ai/benchmark?slug=sciexam-enso-predictivity · How It Works · Data refreshed daily, snapshot 2026-10-09.