Doc2DB-Bench - Entity Tables: leaderboard

Metric: Entity-level cell F1 (%; entity tables extracted from the document; 203 synthesized long-document instances over 42 database schemas from BIRD and Spider in seven domain groups; identical prompts, greedy decoding at temperature 0; cells aligned by global maximum-weight tuple matching, a cell matching on exact numeric equality or at least 90 percent string similarity). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.684.93
2GPT-5.480.09
3Gemini 2.5 Pro78.13
4Qwen 3 Max75.77
5Gemini 2.5 Flash74.21
6GPT-4o70.3
7DeepSeek V4 Flash67.3
8Qwen 2.5 72B Instruct45.03
9Qwen 2.5 14B Instruct44.07
10Llama 3.1 70B Instruct17.25

Interactive version: theaggregate.ai/benchmark?slug=doc2db-bench-entity-tables · How It Works · Data refreshed daily, snapshot 2026-09-29.