DocScope - Fact Extraction: leaderboard

Metric: Fact consistency rate (%): share of extracted facts that a Qwen3.6-Plus judge marks consistent with the gold evidence on the correctly retrieved pages, micro-averaged over the answerable questions of the 730-question DocScope test split (human-annotated questions on long, visually rich PDF documents; the model receives the whole document and returns evidence pages, evidence regions, supporting facts and an answer), each stage scored only on the gold pages the model retrieved; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.777.4
2Gemma 4 31B73.6
3Claude Sonnet 4.672.1
4Gemini 3.1 Flash Lite70.7
5Gemini 3.1 Pro (Preview)68.7
6Qwen 3.6 Plus66.4
7Qwen 3.5 27B61.3
8GPT-5.459.4
9Qwen 3.5 397B A17B58.7
10Qwen 3 VL 235B A22B58.6
11Gemma 4 26B A4B53.1

Interactive version: theaggregate.ai/benchmark?slug=docscope-fact-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.