FutureEval: leaderboard

Metaculus benchmark for AI forecasting accuracy on unresolved real-world questions across science, technology, health, geopolitics, AI, and other domains. Models make probabilistic forecasts and are ranked by a unified log-score-based forecasting score against community and pro forecaster baselines.

Metric: Unified Forecasting Score. Source: www.metaculus.com. Status: saturation imminent. 93 models tracked.

Top models

#ModelScore
1Claude Opus 5 (High)15.41
2Claude Fable 5 (High)13.77
3Claude Opus 4.8 (High)12.97
4Gemini 3.6 Flash12.7
5GPT-5.5 Instant12.66
6GPT-5.6 Sol (High)12.45
7Gemini 3.1 Pro (Preview) (High)12.23
8GPT-5.1 (High)12
9GPT-5.5 (High)11.95
10O311.73
11GPT-511.49
12GPT-5 (High)11.47
13Kimi K311.31
14Gemini 3.5 Flash11.29
15Claude Sonnet 511.18

Interactive version: theaggregate.ai/benchmark?slug=futureeval · How It Works · Data refreshed daily, snapshot 2026-09-05.