SuperCLUE General (November 2025) - Math Reasoning: leaderboard

Metric: Score. Source: www.superclueai.com. 32 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 Speciale75.24
2GPT-5.1 (High)74.07
3DeepSeek V3.2 (Thinking)69.44
4Gemini 3 Pro (Preview)67.59
5O3 (High)66.67
6GPT-5.2 (High)64.49
7Kimi K2 (Thinking)63.89
8GPT-OSS-120B60.65
9Gemini 2.5 Pro60.19
10GLM-4.760
11GLM-4.658.33
12Claude Opus 4.5 (Thinking)58.33
13Grok 4.1 Fast (Reasoning)56.48
14Qwen 3 Max55.56
15Qwen 3 Max (Preview) (Thinking)55.56

Interactive version: theaggregate.ai/benchmark?slug=superclue-general-november-2025-math-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.