HorizonBench - Evolved Preferences: leaderboard

Metric: Accuracy (%) on the 2,484 items whose target preference changed after a life event of HorizonBench (5-option questions on which assistant response fits the user's current preference, given six months of synthetic conversation history, about 163K tokens and 4,300 turns; pre-evolution values as hard distractors; items kept only if five validation models all failed them without the history; chance 20); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 25 models tracked.

Top models

#ModelScore
1Claude Opus 4.551.3
2Claude 3.7 Sonnet46.7
3Claude Opus 444.4
4Gemini 3.1 Pro (Preview)38.4
5Gemini 3 Pro38.4
6Claude Sonnet 4.532.6
7Claude 3.5 Haiku31.8
8Claude Sonnet 431.4
9Claude 3.5 Sonnet (20240620)29.2
10Gemini 2.0 Flash24.2
11GPT-4o23.9
12Gemini 2.5 Pro21.5
13Gemini 2.0 Flash Lite20.3
14Gemini 2.5 Flash Lite18.2
15O115.5

Interactive version: theaggregate.ai/benchmark?slug=horizonbench-evolved-preferences · How It Works · Data refreshed daily, snapshot 2026-10-07.