MarketBench - Self-Assessment Calibration: leaderboard
Metric: Brier skill score (at most 1; 0 matches a constant base-rate forecast, negative is worse) of the model's stated probability that it will solve each of 93 SWE-bench Lite tasks in one attempt, scored against its realized pass or fail; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 | 0.06 |
| 2 | Claude Sonnet 4.5 | 0.02 |
| 3 | Gemini 3 Pro (Preview) | -0.11 |
| 4 | GPT-5.2 Pro | -0.11 |
| 5 | GPT-5.2 | -0.19 |
| 6 | GPT-5 Mini | -0.3 |
Interactive version: theaggregate.ai/benchmark?slug=marketbench-self-assessment-calibration · How It Works · Data refreshed daily, snapshot 2026-10-07.