MarketBench - Self-Assessment Calibration: leaderboard

Metric: Brier skill score (at most 1; 0 matches a constant base-rate forecast, negative is worse) of the model's stated probability that it will solve each of 93 SWE-bench Lite tasks in one attempt, scored against its realized pass or fail; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.50.06
2Claude Sonnet 4.50.02
3Gemini 3 Pro (Preview)-0.11
4GPT-5.2 Pro-0.11
5GPT-5.2-0.19
6GPT-5 Mini-0.3

Interactive version: theaggregate.ai/benchmark?slug=marketbench-self-assessment-calibration · How It Works · Data refreshed daily, snapshot 2026-10-07.