Doc2DB-Bench: leaderboard

Metric: Overall cell-level F1 (%; entity and relationship tables of the constructed database; 203 synthesized long-document instances over 42 database schemas from BIRD and Spider in seven domain groups; identical prompts, greedy decoding at temperature 0; cells aligned by global maximum-weight tuple matching, a cell matching on exact numeric equality or at least 90 percent string similarity). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GPT-5.475.25
2Claude Opus 4.673.6
3Gemini 2.5 Pro70.81
4Gemini 2.5 Flash65.99
5Qwen 3 Max63.44
6GPT-4o59.05
7DeepSeek V4 Flash57.06
8Qwen 2.5 72B Instruct36.07
9Qwen 2.5 14B Instruct32.97
10Llama 3.1 70B Instruct14.29

Interactive version: theaggregate.ai/benchmark?slug=doc2db-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.