ProphetArena: leaderboard

Live forecasting benchmark testing LLMs on real-world prediction tasks from Kalshi markets. Models assign probabilistic estimates; scored by 1-Brier (higher is better). Tests calibration and world knowledge.

Metric: 1 - Brier Score. Source: www.prophetarena.co. Status: saturated. 12 models tracked.

Top models

#ModelScore
1GLM-5.3 Flash0.96
2Gemini 3.7 Flash0.96
3GLM-5.30.96
4DeepSeek V4 Pro (0813)0.96
5Claude Fable 50.95
6Gemini 3.1 Pro (Preview)0.95
7DeepSeek V4 Flash0.95
8Gemini 3 Flash (Preview)0.93
9GPT-5.1 (Medium)0.87
10GPT-5.1 (Low)0.86
11GPT-5.10.86

Interactive version: theaggregate.ai/benchmark?slug=prophetarena · How It Works · Data refreshed daily, snapshot 2026-09-05.