HumRightsBench - Rule Recall: leaderboard
Metric: Accuracy (%; x100 of the exact-set match rate on the rule-recall multiple-choice question: which international human rights rule applies, rules named by instrument and article; pilot HumRightsBench question set on right-to-water scenarios validated by human rights experts; each question asked with five seeds and shuffled answer order (Qwen 3.5-9B multiple-choice types over four), responses through each provider's structured-output interface, errored calls counted as incorrect). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 77.4 |
| 2 | Gemini 3 Flash (Preview) | 76.5 |
| 3 | GPT-5 | 71.5 |
| 4 | Qwen 3.5 9B | 49.4 |
Interactive version: theaggregate.ai/benchmark?slug=humrightsbench-rule-recall · How It Works · Data refreshed daily, snapshot 2026-09-26.