KORGym - Mathematical and Logical — leaderboard

Metric: Score. Source: razor233.github.io. 19 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (03-25)0.93
2O3 Mini0.79
3DeepSeek R10.69
4O1 (2024-12-17)0.65
5Qwen 3 32B (Thinking)0.58
6Claude 3.7 Sonnet (Thinking)0.52
7DeepSeek R1 Distill Qwen 32B0.35
8Gemini 2.0 Flash (Thinking)0.34
9DeepSeek V3 (0324)0.27
10Gemini 2.0 Flash0.17
11Doubao-1.5-Pro0.16
12GPT-4o0.08
13DeepSeek-R1-Distill-Qwen-7B0.06
14Qwen 2.5 72B Instruct0.04
15Qwen 2.5 32B Instruct0.04

Interactive version: theaggregate.ai/benchmark?slug=korgym-mathematical-and-logical · How the rankings work · Data refreshed daily, snapshot 2026-07-22.