NarrativeTrack: leaderboard

Metric: Accuracy (%; mean exact-match over 13 binary, multiple-choice and ordering question cells across five entity-centric reasoning dimensions). Source: arxiv.org. Saturation forecast: Around December 2026. 20 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro83.8
2Gemini 3.1 Pro (Preview)83.4
3Gemini 3 Flash79.32
4GPT-4.174.85
5Gemini 2.5 Flash74.25
6GPT-4o72.27
7Gemini 2.0 Flash60.24
8Qwen 2.5 VL 32B Instruct56.96
9InternVL3-38B55.47
10InternVL3-8B48.81
11Qwen 2.5 VL 7B Instruct48.81

Interactive version: theaggregate.ai/benchmark?slug=narrativetrack · How It Works · Data refreshed daily, snapshot 2026-09-26.