Divergent Thinking — leaderboard

Creativity benchmark where LLMs generate 25 maximally distinct words unrelated to an initial 50-word list. Each word pair is judged by 4 LLMs. Tests originality and fluency.

Metric: Divergence Score. Source: github.com. Status: saturation imminent. 19 models tracked.

Top models

#ModelScore
1O1 Preview4.79
2Gemini 2.0 Flash (Preview)4.65
3Claude 3 Opus4.47
4Grok 2 (1212)4.45
5Llama 3.3 70B4.44
6Claude 3.5 Sonnet (20241022)4.41
7Gemini 2.0 Flash (Thinking)4.41
8Gemma 2 27B4.37
9O1 Mini4.2
10Claude 3.5 Haiku4.16
11Mistral Large 2 (Jul)4.14
12GPT-4o Mini4.12
13Gemini 1.5 Flash4.09
14Gemini 1.5 Pro (Sept)4.07
15Claude 3 Haiku3.98

Interactive version: theaggregate.ai/benchmark?slug=divergent-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.