POLAR-Bench — leaderboard

Metric: Overall Mean (self-reported). Source: benchmarklist.com. 22 models tracked.

Top models

#ModelScore
1GLM-5.194.34
2GPT-5.492.33
3Gemma 4 31B91.2
4Gemma 4 E4B90.87
5DeepSeek V3.190.81
6Llama 3.3 70B Instruct86.23
7Gemma 3 27B83.16
8Qwen 3 32B80.87
9GLM-4.7 Flash80
10Gemma 4 E2B79.74
11Ministral 3 8B76.12
12GPT-OSS-120B74.62
13Ministral 3 14B71.96
14GPT-OSS-20B71.21
15DeepSeek R1 Distill Llama 70B69.93

Interactive version: theaggregate.ai/benchmark?slug=polar-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.