ForecastBench — leaderboard

Measures LLM ability to forecast future events - think prediction markets. 1,000 auto-generated questions from Polymarket, Manifold, and real-world data. Compares against superforecasters and the public.

Metric: Overall Score (higher is better). Source: www.forecastbench.org. Status: saturation imminent. 236 models tracked.

Top models

#ModelScore
1Human Expert70.7
2Grok 4.2068.2
3GPT-566
4Gemini 3 Pro (Preview)65.8
5Grok 465.6
6Claude 3.7 Sonnet (20250219)65.5
7Grok 4 Fast (Reasoning)65.5
8GPT-4.5 (Preview)65.5
9Grok 4.1 Fast (Reasoning)65.2
10Claude Opus 4.1 (20250805)65.2
11Claude Opus 4.5 (20251101)65.2
12Median Human65.1
13Gemini 3.1 Pro (Preview)65
14O3 (2025-04-16)65
15Gemini 3 Flash (Preview)64.9

Interactive version: theaggregate.ai/benchmark?slug=forecastbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.