SynthDocBench - Complex Reasoning: leaderboard

Metric: Accuracy (%) on multi-hop complex-reasoning questions: GPT-5 judge scores each answer 0 to 10 against a deterministic reference and an answer counts when the score is at least 6; pages rasterized at 144 DPI in 5-page strips, vision-only input, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)78.9
2Qwen 3 VL 235B A22B61.1
3GPT-5.445.7
4InternVL3-78B39.7
5GPT-4o36
6Claude Sonnet 4.533.7
7Qwen 2.5 VL 7B Instruct1.2

Interactive version: theaggregate.ai/benchmark?slug=synthdocbench-complex-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.