MCJudgeBench - Intrinsic Inconsistency: leaderboard

Metric: Intrinsic constraint-level inconsistency rate (%, CIR intr): share of constraints whose predicted label is not identical across 5 repeated judgments of the same instance at temperature 1.0; the judge sees an instruction, a candidate response and its explicit constraint list and labels each constraint yes, partial or no, over the 653 constraints of 141 multi-constraint instructions of MCJudgeBench; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Claude Haiku 4.52.94
2Claude Sonnet 4.6 (Thinking)5.21
3Claude Sonnet 4.65.72
4Gemini 3.1 Pro (Preview)6.58
5Gemini 2.5 Flash Lite7.35
6GPT-5.2 (Non-reasoning)9.49
7Claude Haiku 4.5 (Thinking)14.09
8GPT-5.214.7
9Gemini 2.5 Flash Lite (Thinking)14.77
10Qwen 3.5 4B (Non-reasoning)24.2
11Llama 3.2 3B Instruct35.38

Interactive version: theaggregate.ai/benchmark?slug=mcjudgebench-intrinsic-inconsistency · How It Works · Data refreshed daily, snapshot 2026-10-07.