TasteVal: leaderboard

Metric: Compute multiplier against the expert baseline (the best of at least two human expert attempts per task scores 1): GPU hours the baseline needs to reach the final test score of the weaker run divided by the GPU hours the model needs, geometric mean over seeds and then over the 8 withheld AI R&D tasks, chained through GPT-4 and Claude Opus 5 (max) reference runs for models far from expert level; the model only proposes and interprets experiments, which a fixed Claude Opus 4.8 Coder runs on one H100 within 40 H100 hours; each run at the reasoning setting its label names; from 0 with no upper limit; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 24 models tracked.

Top models

#ModelScore
1Claude Opus 5.5 (Max)2.3
2Claude Fable 5.1 (Max)1.72
3GPT-6 Astra (Max)1.6
4Claude Opus 5 (Max)1.29
5Claude Opus 5 (Medium)0.97
6Claude Opus 4.8 (Max)0.84
7GPT-5.6 Sol (Max)0.59
8GLM-5.3 (Max)0.59
9Claude Opus 5 (Low)0.58
10Claude Opus 4.7 (Max)0.5
11GPT-5.5 (xHigh)0.47
12Kimi K3 (Thinking)0.45
13Claude Opus 4.6 (Max)0.4
14GPT-5.2 (High)0.28
15GPT-5.4 (xHigh)0.21

Interactive version: theaggregate.ai/benchmark?slug=tasteval · How It Works · Data refreshed daily, snapshot 2026-10-07.