LongDocBench - Note Relationship Recovery: leaderboard

Metric: Note relationship similarity (%; the model reads fixed TextIn parsing output and layout context of long financial reports, textbooks and papers and links each of 2,680 benchmark-localized tables and figures to its related text; normalized edit similarity to the gold text after IoU matching of objects). Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)42
2Qwen 3.5 397B A17B33
3GLM-5.230
4Qwen 3.5 397B A17B (Non-reasoning)30
5Kimi K2.6 (Non-reasoning)22
6Qwen 3.5 35B A3B (Non-reasoning)20
7MiniMax-M2.517
8Qwen 3.5 9B (Non-reasoning)16

Interactive version: theaggregate.ai/benchmark?slug=longdocbench-note-relationship-recovery · How It Works · Data refreshed daily, snapshot 2026-09-26.