LongDocBench - Contextual Relationship Recovery: leaderboard

Metric: Macro similarity (%; unweighted mean of the caption, note and source relationship scores; the model reads fixed TextIn parsing output and layout context of long financial reports, textbooks and papers and links each of 2,680 benchmark-localized tables and figures to its related text; normalized edit similarity to the gold text after IoU matching of objects). Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)63
2Qwen 3.5 397B A17B60
3Qwen 3.5 397B A17B (Non-reasoning)59
4GLM-5.256
5Kimi K2.6 (Non-reasoning)55
6Qwen 3.5 9B (Non-reasoning)52
7Qwen 3.5 35B A3B (Non-reasoning)52
8MiniMax-M2.544

Interactive version: theaggregate.ai/benchmark?slug=longdocbench-contextual-relationship-recovery · How It Works · Data refreshed daily, snapshot 2026-09-26.