Progress-Bench - Text-Based Demo NSE: leaderboard
Metric: Normalized score error (%; pointwise error of the predicted progress score; answerable items with step-text demonstrations; direct prediction). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 21.1 |
| 2 | GPT-5 | 23.6 |
| 3 | Qwen 3 VL 8B | 35.3 |
| 4 | Qwen 2.5 VL 7B | 39.1 |
Interactive version: theaggregate.ai/benchmark?slug=progress-bench-text-based-demo-nse · How It Works · Data refreshed daily, snapshot 2026-09-26.