JudgeBench Coding — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 52 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)97.62
2GPT-5 Mini97.62
3GPT-5.597.62
4Gemma 4 31B (IT)97.62
5Kimi K2.597.62
6Kimi K2.697.62
7DeepSeek R1 052897.62
8Qwen 3.5 35B A3B97.62
9Step 3.5 Flash97.62
10DeepSeek V3.2 Speciale97.62
11GLM-4.7 FP897.62
12GLM-5 FP897.62
13GLM-5.1 FP897.62
14GPT-OSS-20B96.43
15DeepSeek-V4-Flash-FP896.43

Interactive version: theaggregate.ai/benchmark?slug=judgebench-coding · How the rankings work · Data refreshed daily, snapshot 2026-07-22.