CresOWLve - Exact Match: leaderboard

Metric: Exact-match accuracy (%) after lowercasing and punctuation removal, English version, on the 2,061 creative questions of CresOWLve (expert-written What? Where? When? puzzles that need several real-world facts joined by a creative leap), zero-shot; non-thinking models with chain-of-thought prompting, thinking models at the stated effort; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)51.29
2Gemini 3.1 Pro (Preview) (Medium)49.1
3Gemini 3.1 Pro (Preview) (Low)46.43
4Gemini 3 Flash (Medium)35.08
5Gemini 3 Flash (Minimal)24.99
6Qwen 3.5 397B A17B19.36
7GLM-518.92
8GPT-5.4 (Medium)17.81
9DeepSeek V3.2 (Thinking)15.33
10Qwen 3 235B A22B 2507 (Thinking)13.49
11GPT-4.113.44
12Qwen 3 235B A22B 2507 Instruct13.05
13Mistral-Large-3-675B-Instruct-251211.26
14GPT-5.4 (Non-reasoning)9.41
15GPT-4.1 Mini7.33

Interactive version: theaggregate.ai/benchmark?slug=cresowlve-exact-match · How It Works · Data refreshed daily, snapshot 2026-10-07.