Creative Writing (Lechmazur) — leaderboard
Evaluates creative fiction writing quality across 30+ models using structured prompts and LLM-as-judge scoring. Measures prose quality, creativity, coherence, and adherence to constraints.
Metric: Mean Score. Source: github.com. Status: saturation imminent. 39 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 (High) | 3.3 |
| 2 | GPT-5.5 (xHigh) | 3 |
| 3 | Kimi K3 | 2.9 |
| 4 | GPT-5.6 Sol (xHigh) | 2.9 |
| 5 | GPT-5.4 (xHigh) | 2.7 |
| 6 | GPT-5.6 Sol (High) | 2.7 |
| 7 | GPT-5.4 (Medium) | 2.7 |
| 8 | Claude Opus 4.7 | 2.4 |
| 9 | Claude Opus 4.8 (xHigh) | 1.3 |
| 10 | GPT-5.2 (Medium) | 1 |
| 11 | GLM-5.2 (Max) | 0.9 |
| 12 | Claude Opus 4.8 (High) | 0.8 |
| 13 | Kimi K2.6 | 0.7 |
| 14 | MiniMax-M3 | 0.6 |
| 15 | Mistral Medium 3.1 | 0.2 |
Interactive version: theaggregate.ai/benchmark?slug=creative-writing-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.