TasteVal - Performance Multiplier: leaderboard
Metric: Performance multiplier against the expert baseline: test score of the official submission of each run after the full 40 H100-hour budget, rescaled per task so the naive baseline is 0 and the best human expert attempt is 1 (runs below the naive baseline count 0), arithmetic mean over seeds and then geometric mean over the 8 withheld AI R&D tasks, regardless of compute used; same Researcher and fixed Coder protocol as the compute multiplier; from 0 with no upper limit; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5.5 (Max) | 1.14 |
| 2 | GPT-6 Astra (Max) | 1.13 |
| 3 | Claude Opus 5 (Max) | 1.11 |
| 4 | Claude Fable 5.1 (Max) | 1.09 |
| 5 | Claude Opus 5 (Medium) | 1.07 |
| 6 | Claude Opus 4.8 (Max) | 1.01 |
| 7 | GPT-5.6 Sol (Max) | 0.92 |
| 8 | Claude Opus 5 (Low) | 0.92 |
| 9 | GLM-5.3 (Max) | 0.89 |
| 10 | Kimi K3 (Thinking) | 0.89 |
| 11 | GPT-5.5 (xHigh) | 0.84 |
| 12 | Claude Opus 4.6 (Max) | 0.78 |
| 13 | Claude Opus 4.7 (Max) | 0.75 |
| 14 | GPT-5.2 (High) | 0.7 |
| 15 | GPT-5.4 (xHigh) | 0.6 |
Interactive version: theaggregate.ai/benchmark?slug=tasteval-performance-multiplier · How It Works · Data refreshed daily, snapshot 2026-10-07.