MCJudgeBench - Macro-F1: leaderboard

Metric: Macro-F1 (%, times 100) over the yes, partial and no classes of constraint-level verdicts: the judge sees an instruction, a candidate response and its explicit constraint list and labels each constraint yes, partial or no, over the 653 constraints of 141 multi-constraint instructions of MCJudgeBench; deterministic decoding at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 11 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Thinking)63.7
2GPT-5.261.6
3Claude Sonnet 4.660.7
4Claude Haiku 4.5 (Thinking)60.3
5Gemini 3.1 Pro (Preview)59.8
6GPT-5.2 (Non-reasoning)59.2
7Claude Haiku 4.558.5
8Gemini 2.5 Flash Lite (Thinking)56.9
9Qwen 3.5 4B (Non-reasoning)52.9
10Gemini 2.5 Flash Lite51.8
11Llama 3.2 3B Instruct44.1

Interactive version: theaggregate.ai/benchmark?slug=mcjudgebench-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.