TasteVal: leaderboard
Metric: Compute multiplier against the expert baseline (the best of at least two human expert attempts per task scores 1): GPU hours the baseline needs to reach the final test score of the weaker run divided by the GPU hours the model needs, geometric mean over seeds and then over the 8 withheld AI R&D tasks, chained through GPT-4 and Claude Opus 5 (max) reference runs for models far from expert level; the model only proposes and interprets experiments, which a fixed Claude Opus 4.8 Coder runs on one H100 within 40 H100 hours; each run at the reasoning setting its label names; from 0 with no upper limit; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5.5 (Max) | 2.3 |
| 2 | Claude Fable 5.1 (Max) | 1.72 |
| 3 | GPT-6 Astra (Max) | 1.6 |
| 4 | Claude Opus 5 (Max) | 1.29 |
| 5 | Claude Opus 5 (Medium) | 0.97 |
| 6 | Claude Opus 4.8 (Max) | 0.84 |
| 7 | GPT-5.6 Sol (Max) | 0.59 |
| 8 | GLM-5.3 (Max) | 0.59 |
| 9 | Claude Opus 5 (Low) | 0.58 |
| 10 | Claude Opus 4.7 (Max) | 0.5 |
| 11 | GPT-5.5 (xHigh) | 0.47 |
| 12 | Kimi K3 (Thinking) | 0.45 |
| 13 | Claude Opus 4.6 (Max) | 0.4 |
| 14 | GPT-5.2 (High) | 0.28 |
| 15 | GPT-5.4 (xHigh) | 0.21 |
Interactive version: theaggregate.ai/benchmark?slug=tasteval · How It Works · Data refreshed daily, snapshot 2026-10-07.