CulMind-R: leaderboard

Metric: Final-answer score (0-100) on the 24-task CulMind-R reasoning subset, the model writing a structured reasoning process before its answer; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct64.9
2GPT-5.563.5
3Gemini 3 Flash (Preview)61.3
4Qwen 3.5 27B59.6
5GLM-4.5V58.3
6GPT-5.4 Mini54.6
7Qwen 3 VL 8B Instruct54.4
8Qwen 3.5 9B53.5
9InternVL3-14B51.3
10InternVL3-8B50.8
11Gemini 2.5 Flash48.2

Interactive version: theaggregate.ai/benchmark?slug=culmind-r · How It Works · Data refreshed daily, snapshot 2026-09-29.