Fiction.LiveBench — leaderboard

Long-context comprehension benchmark using 36 questions across 30 stories at varying lengths. Tests deep narrative understanding, subtext, and plot hole detection - not just retrieval.

Metric: Accuracy (%). Source: fiction.live. Status: saturated. 41 models tracked.

Top models

#ModelScore
1O3 (2025-04-16) (Medium)100
2Grok 496.9
3GPT-5 (Medium)96.9
4Gemini 2.5 Pro (03-25)90.6
5Gemini 2.5 Pro (Preview 06-05)87.5
6Kimi K2.578.1
7Grok 4 Fast75
8Gemini 2.5 Pro (Preview 03-25)71.9
9Gemini 2.5 Pro (Preview 05-06)71.9
10Qwen 3 235B A22B (Thinking)68.8
11Gemini 2.5 Flash (Preview 05-20)68.8
12GPT-4.5 (Preview)63.9
13GPT-4.162.5
14Qwen 3 Next 80B A3B Instruct62.5
15Gemini 2.0 Flash (001)62.5

Interactive version: theaggregate.ai/benchmark?slug=fiction-livebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.