Story Theory Bench — leaderboard

Creative narrative generation benchmark: LLMs write short stories incorporating 10 required elements (characters, objects, concepts, attributes, motivations). Scored 50% programmatic + 50% LLM judge across 34 tasks.

Metric: Score (%). Source: github.com. Status: saturated. 35 models tracked.

Top models

#ModelScore
1GPT-5.499.6
2GLM-599.6
3Mercury 299.1
4Qwen 3.5 27B98.9
5Gemini 3.1 Flash Lite (Preview)98.4
6Claude Opus 4.698.1
7DeepSeek V3.292.2
8Claude Opus 4.590.9
9Claude Sonnet 4.590.1
10Step 3.5 Flash90.1
11Claude Sonnet 489.6
12O389.5
13GLM-4.788.8
14Kimi K2 (Thinking)88.7
15Gemini 3 Flash (Preview)88.3

Interactive version: theaggregate.ai/benchmark?slug=story-theory-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.