RubricBench (Self-Generated Rubrics): leaderboard

Metric: Preference accuracy (%) on RubricBench's 1,147 preference pairs (instruction following, STEM, code, safety and chat pairs re-curated from HelpSteer3, PPE and RewardBench2 and filtered so that surface cues such as length, formatting or tone favour the rejected response); the judge first writes a rubric of binary checks from the instruction and then verifies both responses against it (the OpenRubric pipeline of Table 2); higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 7 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro60.4#77
2Qwen 3.5 Plus59.3#123
3Gemini 3 Flash58#93
4DeepSeek V3.2 (Non-reasoning)57.8#198 (DeepSeek V3.2)
5GPT-OSS-120B56.4#330
6GPT-5.154.6#131
7GPT-4o Mini46.7#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=rubricbench-self-generated-rubrics · How It Works · Data refreshed daily, snapshot 2026-10-11.