RINoBench - Justification Alignment: leaderboard

Metric: Justification alignment (ALI, 0-1 G-Eval score from an LLM judge, shown times 100): whether the model's novelty justification follows the same arguments and conclusion as the gold justification synthesized from the human reviews, on the RINoBench test split (one fifth of the 1,381 ICLR 2022-2023 research ideas, gold scores from human reviewers), zero-shot with the five-point novelty rubric and the idea's related works; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 7 models tracked.

Top models

#ModelScoreOverall rank
1O372#121
2GPT-571#91
3DeepSeek R167#245
4GPT-OSS-120B64#330
5Llama 3.1 8B58#1139
6Llama 4 Scout58#646
7Llama 3.3 70B Instruct55#520

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=rinobench-justification-alignment · How It Works · Data refreshed daily, snapshot 2026-10-11.