ConStory-Bench — leaderboard

Long story generation benchmark measuring cross-scene consistency bugs using Consistency Error Density over 2,000 generated stories.

Metric: CED errors per 10K words (lower is better). Source: picrew.github.io. Status: saturation imminent. 66 models tracked.

Top models

#ModelScore
1GPT-50.11
2Gemini 2.5 Pro0.3
3Gemini 2.5 Flash0.31
4Claude Sonnet 4.50.52
5GLM-4.60.53
6Qwen 3 32B0.54
7Ring-1T0.54
8DeepSeek V3.2 Exp0.54
9Qwen 3 235B A22B (Thinking)0.56
10GLM-4.50.6
11Grok 40.67
12Ling-1T0.7
13Qwen 3 Next 80B A3B (Thinking)0.96
14Kimi K2 (0711)1.33
15Mistral Medium 3.11.36

Interactive version: theaggregate.ai/benchmark?slug=constory-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.