Fiction.LiveBench — leaderboard
Long-context comprehension benchmark using 36 questions across 30 stories at varying lengths. Tests deep narrative understanding, subtext, and plot hole detection - not just retrieval.
Metric: Accuracy (%). Source: fiction.live. Status: saturated. 41 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 (2025-04-16) (Medium) | 100 |
| 2 | Grok 4 | 96.9 |
| 3 | GPT-5 (Medium) | 96.9 |
| 4 | Gemini 2.5 Pro (03-25) | 90.6 |
| 5 | Gemini 2.5 Pro (Preview 06-05) | 87.5 |
| 6 | Kimi K2.5 | 78.1 |
| 7 | Grok 4 Fast | 75 |
| 8 | Gemini 2.5 Pro (Preview 03-25) | 71.9 |
| 9 | Gemini 2.5 Pro (Preview 05-06) | 71.9 |
| 10 | Qwen 3 235B A22B (Thinking) | 68.8 |
| 11 | Gemini 2.5 Flash (Preview 05-20) | 68.8 |
| 12 | GPT-4.5 (Preview) | 63.9 |
| 13 | GPT-4.1 | 62.5 |
| 14 | Qwen 3 Next 80B A3B Instruct | 62.5 |
| 15 | Gemini 2.0 Flash (001) | 62.5 |
Interactive version: theaggregate.ai/benchmark?slug=fiction-livebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.