ConStory-Bench — leaderboard
Long story generation benchmark measuring cross-scene consistency bugs using Consistency Error Density over 2,000 generated stories.
Metric: CED errors per 10K words (lower is better). Source: picrew.github.io. Status: saturation imminent. 66 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 0.11 |
| 2 | Gemini 2.5 Pro | 0.3 |
| 3 | Gemini 2.5 Flash | 0.31 |
| 4 | Claude Sonnet 4.5 | 0.52 |
| 5 | GLM-4.6 | 0.53 |
| 6 | Qwen 3 32B | 0.54 |
| 7 | Ring-1T | 0.54 |
| 8 | DeepSeek V3.2 Exp | 0.54 |
| 9 | Qwen 3 235B A22B (Thinking) | 0.56 |
| 10 | GLM-4.5 | 0.6 |
| 11 | Grok 4 | 0.67 |
| 12 | Ling-1T | 0.7 |
| 13 | Qwen 3 Next 80B A3B (Thinking) | 0.96 |
| 14 | Kimi K2 (0711) | 1.33 |
| 15 | Mistral Medium 3.1 | 1.36 |
Interactive version: theaggregate.ai/benchmark?slug=constory-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.