HorizonBench: leaderboard

Metric: Accuracy (%) on all 4,245 items (360 users, three generator models) of HorizonBench (5-option questions on which assistant response fits the user's current preference, given six months of synthetic conversation history, about 163K tokens and 4,300 turns; pre-evolution values as hard distractors; items kept only if five validation models all failed them without the history; chance 20); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 25 models tracked.

Top models

#ModelScore
1Claude Opus 4.552.8
2Claude 3.7 Sonnet47.8
3Claude Opus 445.1
4Gemini 3.1 Pro (Preview)41.2
5Gemini 3 Pro40.9
6Claude Sonnet 4.534
7Claude 3.5 Haiku32.5
8Claude Sonnet 432.2
9Claude 3.5 Sonnet (20240620)29.8
10Gemini 2.0 Flash24.7
11GPT-4o24.6
12Gemini 2.5 Pro22.8
13Gemini 2.0 Flash Lite20.9
14Gemini 2.5 Flash Lite19.5
15O316.5

Interactive version: theaggregate.ai/benchmark?slug=horizonbench · How It Works · Data refreshed daily, snapshot 2026-10-07.