LongDocBench - Source Relationship Recovery: leaderboard

Metric: Source relationship similarity (%; the model reads fixed TextIn parsing output and layout context of long financial reports, textbooks and papers and links each of 2,680 benchmark-localized tables and figures to its related text; normalized edit similarity to the gold text after IoU matching of objects). Source: arxiv.org. Saturation forecast: Around April 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B (Non-reasoning)85
2Qwen 3.5 397B A17B82
3Kimi K2.6 (Non-reasoning)82
4GPT-5.6 Sol (xHigh)81
5Qwen 3.5 9B (Non-reasoning)80
6Qwen 3.5 35B A3B (Non-reasoning)78
7GLM-5.274
8MiniMax-M2.561

Interactive version: theaggregate.ai/benchmark?slug=longdocbench-source-relationship-recovery · How It Works · Data refreshed daily, snapshot 2026-09-26.