DataGovBench - Table QA: leaderboard

Metric: Question-set accuracy (%; 211 simple and decomposable questions over 178 government open-data datasets, a set counts only when every sub-question is right; text answers by exact match, charts by a majority of four MLLM judges; the model writes Python once from the first 10 table rows, without the Answer Agent). Source: arxiv.org. Saturation forecast: Around March 2027. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.633.7
2Gemini 2.5 Flash31
3GPT-5.128.9
4GPT-4o (2024-08-06)24.2
5Qwen 3 30B A3B 2507 Instruct13.4
6Qwen 3 Coder 30B A3B Instruct12.3
7DeepSeek R1 Distill Qwen 14B3.8
8Llama 3.1 8B Instruct1.9

Interactive version: theaggregate.ai/benchmark?slug=datagovbench-table-qa · How It Works · Data refreshed daily, snapshot 2026-09-29.