TasteVal - Performance Multiplier: leaderboard

Metric: Performance multiplier against the expert baseline: test score of the official submission of each run after the full 40 H100-hour budget, rescaled per task so the naive baseline is 0 and the best human expert attempt is 1 (runs below the naive baseline count 0), arithmetic mean over seeds and then geometric mean over the 8 withheld AI R&D tasks, regardless of compute used; same Researcher and fixed Coder protocol as the compute multiplier; from 0 with no upper limit; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 24 models tracked.

Top models

#ModelScore
1Claude Opus 5.5 (Max)1.14
2GPT-6 Astra (Max)1.13
3Claude Opus 5 (Max)1.11
4Claude Fable 5.1 (Max)1.09
5Claude Opus 5 (Medium)1.07
6Claude Opus 4.8 (Max)1.01
7GPT-5.6 Sol (Max)0.92
8Claude Opus 5 (Low)0.92
9GLM-5.3 (Max)0.89
10Kimi K3 (Thinking)0.89
11GPT-5.5 (xHigh)0.84
12Claude Opus 4.6 (Max)0.78
13Claude Opus 4.7 (Max)0.75
14GPT-5.2 (High)0.7
15GPT-5.4 (xHigh)0.6

Interactive version: theaggregate.ai/benchmark?slug=tasteval-performance-multiplier · How It Works · Data refreshed daily, snapshot 2026-10-07.