NARU (Open-Ended): leaderboard

Metric: Atomic-fact recall (%; x100 of the 0-1 FActScore-style recall: share of the reference answer's atomic facts that the free-form answer covers, judged by GPT-5.5, mean of the nine categories over a stratified 500-question subset of NARU with the answer options removed). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Gemini 3 Flash78
2Gemini 3 Pro75
3Gemini 2.5 Flash66
4Qwen 3.5 9B56
5Qwen 3 VL 8B44
6Qwen 2.5 VL 7B35

Interactive version: theaggregate.ai/benchmark?slug=naru-open-ended · How It Works · Data refreshed daily, snapshot 2026-09-26.