Forecast-Dojo (Research Tools + Memory): leaderboard
Metric: Multiclass Brier score (0-2 scale, lower is better; mean over 797 scheduled forecast dates of 230 held-out Polymarket events resolved March-June 2026, four rollouts each; unusable forecasts scored as a uniform distribution; the same research tools plus the agent's own belief notebook carried from the previous forecast date). Source: arxiv.org. Saturation forecast: Around 2030. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Max) | 0.55 |
| 2 | GPT-5.5 (xHigh) | 0.57 |
| 3 | GPT-5.4 (xHigh) | 0.58 |
| 4 | Claude Opus 4.6 (Max) | 0.6 |
| 5 | Claude Opus 4.8 (Max) | 0.61 |
| 6 | MiniMax-M2.5 | 0.64 |
| 7 | Qwen 3.5 397B A17B | 0.64 |
| 8 | DeepSeek V3.2 (High) | 0.67 |
| 9 | GPT-OSS-120B (High) | 0.7 |
Interactive version: theaggregate.ai/benchmark?slug=forecast-dojo-research-tools-plus-memory · How It Works · Data refreshed daily, snapshot 2026-09-26.