Divergent Thinking: leaderboard

Creativity benchmark where LLMs generate 25 maximally distinct words unrelated to an initial 50-word list. Each word pair is judged by 4 LLMs. Tests originality and fluency.

Metric: Divergence Score. Source: github.com. Status: saturated. 19 models tracked.

Top models

#ModelScore
1O1 Preview4.79
2Claude 3 Opus4.47
3Grok 2 (1212)4.45
4Llama 3.3 70B4.44
5Claude 3.5 Sonnet (20241022)4.41
6Gemini 2.0 Flash (Thinking)4.41
7Gemma 2 27B4.37
8O1 Mini4.2
9Claude 3.5 Haiku4.16
10Mistral Large 2 (Jul)4.14
11GPT-4o Mini4.12
12Gemini 1.5 Flash4.09
13Gemini 1.5 Pro (Sept)4.07
14Claude 3 Haiku3.98
15Qwen 2.5 72B3.89

Interactive version: theaggregate.ai/benchmark?slug=divergent-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.