CreativityBench: leaderboard

Metric: Gold correct rate (%, times 100): share of tasks where the model selects both the gold entity and the gold part, exact match, over the 14K affordance-based creative tool-use tasks of CreativityBench (eight household scenes; the model sees the task and every entity in the scene with its parts and attributes, and must name the entity and part to repurpose and how to use it), temperature 0 except GPT-5 Mini and Nano (fixed default sampling); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScore
1Qwen 3 32B25.88
2Qwen 3 14B24.83
3Llama 3.3 70B Instruct21.51
4Ministral-3-14B-Reasoning-251220.91
5Qwen 3 4B18.82
6GPT-5.218.19
7GPT-5 Mini16.87
8Gemini 2.5 Pro16.7
9Gemini 2.5 Flash15.32
10GPT-5 Nano11.92

Interactive version: theaggregate.ai/benchmark?slug=creativitybench · How It Works · Data refreshed daily, snapshot 2026-10-07.