Creative Writing (Lechmazur): leaderboard
Evaluates creative fiction writing quality across 30+ models using structured prompts and LLM-as-judge scoring. Measures prose quality, creativity, coherence, and adherence to constraints.
Metric: Mean Score. Source: github.com. Status: years away from saturation. 48 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (xHigh) | 4.1 |
| 2 | Claude Fable 5.1 (High) | 4.1 |
| 3 | GLM-5.3 (Max) | 3.4 |
| 4 | Claude Fable 5 (High) | 3.1 |
| 5 | Kimi K3 | 2.8 |
| 6 | GPT-5.5 (xHigh) | 2.7 |
| 7 | GPT-5.6 Sol (xHigh) | 2.7 |
| 8 | GPT-5.6 Sol (High) | 2.6 |
| 9 | GPT-5.4 (xHigh) | 2.2 |
| 10 | GPT-5.4 (Medium) | 2.1 |
| 11 | Claude Opus 4.7 | 2 |
| 12 | Claude Opus 4.8 (xHigh) | 1.1 |
| 13 | Muse Spark 1.1 (High) | 1.1 |
| 14 | DeepSeek V4 Pro (High) | 0.8 |
| 15 | GLM-5.2 (Max) | 0.7 |
Interactive version: theaggregate.ai/benchmark?slug=creative-writing-lechmazur · How It Works · Data refreshed daily, snapshot 2026-09-05.