CresOWLve - Russian Exact Match: leaderboard

Metric: Exact-match accuracy (%) after lowercasing, punctuation removal and Unicode normalization, original Russian version, on the 2,061 creative questions of CresOWLve (expert-written What? Where? When? puzzles that need several real-world facts joined by a creative leap), zero-shot; non-thinking models with chain-of-thought prompting, thinking models at the stated effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)67.78
2Gemini 3.1 Pro (Preview) (Medium)65.79
3Gemini 3.1 Pro (Preview) (Low)64.1
4Gemini 3 Flash (Medium)51.82
5Gemini 3 Flash (Minimal)39.98
6GLM-528.87
7GPT-5.4 (Medium)26.88
8Qwen 3.5 397B A17B26.15
9GPT-4.117.03
10Mistral-Large-3-675B-Instruct-251215.48
11GPT-5.4 (Non-reasoning)15.38
12Qwen 3 235B A22B 2507 Instruct13.54
13Qwen 3 235B A22B 2507 (Thinking)13.34
14DeepSeek V3.2 (Thinking)10.48
15GPT-4.1 Mini4.85

Interactive version: theaggregate.ai/benchmark?slug=cresowlve-russian-exact-match · How It Works · Data refreshed daily, snapshot 2026-10-07.