The Aggregate - Forecasting Resolved Questions: leaderboard
Metric: Brier score. Source: theaggregate.ai. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Opus 5 (High) | 0.27 | #11 (Claude Opus 5) |
| 2 | GPT-5.4 (High) | 0.29 | #76 (GPT-5.4) |
| 3 | Claude Opus 4.8 (High) | 0.3 | #33 (Claude Opus 4.8) |
| 4 | GPT-5.5 (High) | 0.31 | #26 (GPT-5.5) |
| 5 | Claude Fable 5 (High) | 0.31 | #13 (Claude Fable 5) |
| 6 | GPT-5.2 (High) | 0.33 | #105 (GPT-5.2) |
| 7 | DeepSeek V4 Pro (High) | 0.33 | #96 (DeepSeek V4 Pro) |
| 8 | Kimi K2.6 (High) | 0.34 | #99 (Kimi K2.6) |
| 9 | Qwen 3.6 Plus (High) | 0.36 | #113 (Qwen 3.6 Plus) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=the-aggregate-forecasting-resolved-questions · How It Works · Data refreshed daily, snapshot 2026-10-11.