ForecastBench: leaderboard

Measures LLM ability to forecast future events - think prediction markets. 1,000 auto-generated questions from Polymarket, Manifold, and real-world data. Compares against superforecasters and the public.

Metric: Overall Score (higher is better). Source: www.forecastbench.org. Status: years away from saturation. 299 models tracked.

Top models

#ModelScore
1Median Expert70.7
2Grok 4.2067.9
3GPT-565.7
4Gemini 3 Pro (Preview)65.6
5Grok 465.3
6Grok 4 Fast (Reasoning)65.3
7Claude 3.7 Sonnet (20250219)65.2
8GPT-4.5 (Preview)65.2
9Grok 4.1 Fast (Reasoning)65.1
10Claude Opus 4.1 (20250805)65.1
11Claude Opus 4.5 (20251101)65.1
12Median Human65.1
13Gemini 3.1 Pro (Preview)64.9
14Gemini 3 Flash (Preview)64.9
15O3 (2025-04-16)64.9

Interactive version: theaggregate.ai/benchmark?slug=forecastbench · How It Works · Data refreshed daily, snapshot 2026-09-05.