FormStruct-Bench - Path Accuracy: leaderboard

Metric: TSR-path (%; share of ground-truth leaf field-path and value pairs reproduced exactly, a strict leaf recall; macro-averaged over the 1,100 human-reviewed pages of the template-disjoint test split of table-form documents in seven languages; the model outputs the page as FormStruct-Bench JSON (deterministic decoding where available; pipelines as released), and missing, unparsable or schema-invalid outputs score 0). Source: arxiv.org. Saturation forecast: Around 2030. 14 models tracked.

Top models

#ModelScore
1Qwen 3.6 35B A3B17.91
2GPT-5.517.15
3Seed 2.1 Pro15.9
4Gemini 3.5 Flash14.55
5Claude Sonnet 513.77
6Kimi K2.612.96
7Qwen 3.5 9B12.58
8Qwen 3.7 Plus9.75
9Step3 VL 10B0.04

Interactive version: theaggregate.ai/benchmark?slug=formstruct-bench-path-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.