ResearchQA (Open Paper) - Refusal Correctness: leaderboard

Metric: Refusal correctness (%; share of the adversarial false-premise questions in the sample answered with a refusal or with only grounded citations, citation-grounded answers from the Open Paper chat-with-paper harness to a 100-row evenly spaced sample of 6,211 single-paper questions over 494 open-access papers, citations checked by a normalizing substring matcher against the extracted paper text; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)87
2GPT-4.187
3GPT-5.478.3
4GLM-4.778.3
5Claude Opus 4.773.9
6Claude Haiku 4.573.9
7Gemini 3 Flash (Preview)65.2
8GPT-OSS-120B47.8

Interactive version: theaggregate.ai/benchmark?slug=researchqa-open-paper-refusal-correctness · How It Works · Data refreshed daily, snapshot 2026-09-29.