HumRightsBench - Rule Application: leaderboard

Metric: Accuracy (%; x100 of the share of rule-application answers whose ranking of 5-7 applicable rules by authority and relevance reaches Kendall tau of at least 0.7 against the expert ranking, unparseable rankings scored 0; pilot HumRightsBench question set on right-to-water scenarios validated by human rights experts; each question asked with five seeds and shuffled answer order (Qwen 3.5-9B multiple-choice types over four), responses through each provider's structured-output interface, errored calls counted as incorrect). Source: arxiv.org. Saturation forecast: Around December 2027. 4 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)24
2GPT-522.5
3Claude Opus 4.718
4Qwen 3.5 9B2.5

Interactive version: theaggregate.ai/benchmark?slug=humrightsbench-rule-application · How It Works · Data refreshed daily, snapshot 2026-09-26.