CresOWLve - Russian: leaderboard

Metric: Accuracy (%) judged by GPT-4o against the reference answer, its explanation and accepted alternatives, original Russian version, on the 2,061 creative questions of CresOWLve (expert-written What? Where? When? puzzles that need several real-world facts joined by a creative leap), zero-shot; non-thinking models with chain-of-thought prompting, thinking models at the stated effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)85.74
2Gemini 3.1 Pro (Preview) (Medium)83.6
3Gemini 3.1 Pro (Preview) (Low)80.49
4GPT-5.4 (Medium)68.22
5Gemini 3 Flash (Medium)67.25
6Gemini 3 Flash (Minimal)54.34
7GLM-544.25
8Qwen 3.5 397B A17B35.71
9GPT-5.4 (Non-reasoning)26.54
10DeepSeek V3.2 (Thinking)26.44
11GPT-4.126.06
12Mistral-Large-3-675B-Instruct-251222.22
13Qwen 3 235B A22B 2507 (Thinking)21.11
14Qwen 3 235B A22B 2507 Instruct20.14
15GPT-4.1 Mini8.83

Interactive version: theaggregate.ai/benchmark?slug=cresowlve-russian · How It Works · Data refreshed daily, snapshot 2026-10-07.