NARU (Open-Ended): leaderboard
Metric: Atomic-fact recall (%; x100 of the 0-1 FActScore-style recall: share of the reference answer's atomic facts that the free-form answer covers, judged by GPT-5.5, mean of the nine categories over a stratified 500-question subset of NARU with the answer options removed). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 78 |
| 2 | Gemini 3 Pro | 75 |
| 3 | Gemini 2.5 Flash | 66 |
| 4 | Qwen 3.5 9B | 56 |
| 5 | Qwen 3 VL 8B | 44 |
| 6 | Qwen 2.5 VL 7B | 35 |
Interactive version: theaggregate.ai/benchmark?slug=naru-open-ended · How It Works · Data refreshed daily, snapshot 2026-09-26.