CresOWLve: leaderboard

Metric: Accuracy (%) judged by GPT-4o against the reference answer, its explanation and accepted alternatives, English version, on the 2,061 creative questions of CresOWLve (expert-written What? Where? When? puzzles that need several real-world facts joined by a creative leap), zero-shot; non-thinking models with chain-of-thought prompting, thinking models at the stated effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)76.03
2Gemini 3.1 Pro (Preview) (Medium)72.97
3Gemini 3.1 Pro (Preview) (Low)67.54
4GPT-5.4 (Medium)53.47
5Gemini 3 Flash (Medium)52.16
6Gemini 3 Flash (Minimal)41.1
7GLM-537.17
8Qwen 3.5 397B A17B31.25
9DeepSeek V3.2 (Thinking)25.47
10GPT-4.124.99
11GPT-5.4 (Non-reasoning)20.96
12Qwen 3 235B A22B 2507 Instruct20.82
13Qwen 3 235B A22B 2507 (Thinking)20.82
14Mistral-Large-3-675B-Instruct-251217.61
15GPT-4.1 Mini13.78

Interactive version: theaggregate.ai/benchmark?slug=cresowlve · How It Works · Data refreshed daily, snapshot 2026-10-07.