FormStruct-Bench - Schema Similarity: leaderboard

Metric: Schema-nTED (%; 1 minus the normalized tree edit distance between predicted and ground-truth schema trees that keep field names, container types and hierarchy but drop values; macro-averaged over the 1,100 human-reviewed pages of the template-disjoint test split of table-form documents in seven languages; the model outputs the page as FormStruct-Bench JSON (deterministic decoding where available; pipelines as released), and missing, unparsable or schema-invalid outputs score 0). Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1GPT-5.554.15
2Seed 2.1 Pro52.81
3Claude Sonnet 552.22
4Gemini 3.5 Flash48.59
5Qwen 3.6 35B A3B48.46
6Kimi K2.648.2
7Qwen 3.5 9B45.58
8Qwen 3.7 Plus34.01
9Step3 VL 10B26.2

Interactive version: theaggregate.ai/benchmark?slug=formstruct-bench-schema-similarity · How It Works · Data refreshed daily, snapshot 2026-09-26.