ProphetArena — leaderboard

Live forecasting benchmark testing LLMs on real-world prediction tasks from Kalshi markets. Models assign probabilistic estimates; scored by 1-Brier (higher is better). Tests calibration and world knowledge.

Metric: 1 - Brier Score. Source: www.prophetarena.co. Status: saturated. 50 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.95
2GPT-5.5 (High)0.95
3Claude Opus 4.8 (Thinking)0.95
4Gemini 3.5 Flash0.95
5Grok 4.200.95
6GLM-5.10.94
7Claude Sonnet 4.60.94
8DeepSeek V4 Pro0.94
9Kimi K2.50.94
10Grok 4.30.94
11MiniMax-M2.70.94
12Qwen 3.6 Plus0.94
13Kimi K2.60.94
14Gemini 3 Flash (Preview)0.93
15Gemini 3 Pro (Preview)0.91

Interactive version: theaggregate.ai/benchmark?slug=prophetarena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.