ProphetArena: leaderboard
Live forecasting benchmark testing LLMs on real-world prediction tasks from Kalshi markets. Models assign probabilistic estimates; scored by 1-Brier (higher is better). Tests calibration and world knowledge.
Metric: 1 - Brier Score. Source: www.prophetarena.co. Status: saturated. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.3 Flash | 0.96 |
| 2 | Gemini 3.7 Flash | 0.96 |
| 3 | GLM-5.3 | 0.96 |
| 4 | DeepSeek V4 Pro (0813) | 0.96 |
| 5 | Claude Fable 5 | 0.95 |
| 6 | Gemini 3.1 Pro (Preview) | 0.95 |
| 7 | DeepSeek V4 Flash | 0.95 |
| 8 | Gemini 3 Flash (Preview) | 0.93 |
| 9 | GPT-5.1 (Medium) | 0.87 |
| 10 | GPT-5.1 (Low) | 0.86 |
| 11 | GPT-5.1 | 0.86 |
Interactive version: theaggregate.ai/benchmark?slug=prophetarena · How It Works · Data refreshed daily, snapshot 2026-09-05.