COHERENCE - StoryBird: leaderboard

Metric: Exact-match accuracy (%) on the 1,398 StoryBird story documents; 6,161 interleaved image-text documents whose images were removed: the model assigns each candidate image to its placeholder; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)63.59
2GPT-5.4 (High)57.37
3Qwen 3.5 397B A17B55.58
4Claude Sonnet 4.6 (Thinking)52.72
5Doubao-Seed-2.0-Pro-26021552.22
6Qwen 3.5 122B A10B49.93
7Qwen 3.5 35B A3B48.21
8Kimi K2.546.21
9GPT-5.237.34
10Qwen 3.5 4B36.41
11Qwen 3 VL 235B A22B32.12
12GLM-4.6V29.18
13Step3 VL 10B25.25
14Qwen 3 VL 8B Instruct24.18
15Qwen 3 VL 4B Instruct22.46

Interactive version: theaggregate.ai/benchmark?slug=coherence-storybird · How It Works · Data refreshed daily, snapshot 2026-10-07.