FutureEval — leaderboard

Metaculus benchmark for AI forecasting accuracy on unresolved real-world questions across science, technology, health, geopolitics, AI, and other domains. Models make probabilistic forecasts and are ranked by a unified log-score-based forecasting score against community and pro forecaster baselines.

Metric: Unified Forecasting Score. Source: www.metaculus.com. Status: saturation imminent. 38 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)19.84
2Gemini 3.1 Pro (Preview)19.03
3Grok 4.20 Multi-Agent14.99
4Claude Opus 4.7 (High)14.62
5GPT-5.5 (High)14.06
6GPT-5.5 Instant13.84
7GPT-5.1 (High)12.14
8O311.94
9GPT-511.49
10GPT-5 (High)11.4
11Kimi K2.611.39
12Grok 4.3 (High)11.35
13GPT-5.4 (High)11.2
14Gemini 3 Pro10.57
15GPT-5.2 (High)10.55

Interactive version: theaggregate.ai/benchmark?slug=futureeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.