FormStruct-Bench - Value Similarity: leaderboard

Metric: Value-nED (%; normalized Levenshtein similarity of non-empty leaf values under maximum-weight one-to-one matching, divided by the larger value count; macro-averaged over the 1,100 human-reviewed pages of the template-disjoint test split of table-form documents in seven languages; the model outputs the page as FormStruct-Bench JSON (deterministic decoding where available; pipelines as released), and missing, unparsable or schema-invalid outputs score 0). Source: arxiv.org. Saturation forecast: Around March 2027. 14 models tracked.

Top models

#ModelScore
1GPT-5.580.98
2Seed 2.1 Pro77.45
3Kimi K2.676.63
4Claude Sonnet 575.42
5Qwen 3.7 Plus74.02
6Qwen 3.6 35B A3B66.93
7Gemini 3.5 Flash66.8
8Qwen 3.5 9B65.73
9Step3 VL 10B59.48

Interactive version: theaggregate.ai/benchmark?slug=formstruct-bench-value-similarity · How It Works · Data refreshed daily, snapshot 2026-09-26.