DataGovBench - Table QA: leaderboard
Metric: Question-set accuracy (%; 211 simple and decomposable questions over 178 government open-data datasets, a set counts only when every sub-question is right; text answers by exact match, charts by a majority of four MLLM judges; the model writes Python once from the first 10 table rows, without the Answer Agent). Source: arxiv.org. Saturation forecast: Around March 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 33.7 |
| 2 | Gemini 2.5 Flash | 31 |
| 3 | GPT-5.1 | 28.9 |
| 4 | GPT-4o (2024-08-06) | 24.2 |
| 5 | Qwen 3 30B A3B 2507 Instruct | 13.4 |
| 6 | Qwen 3 Coder 30B A3B Instruct | 12.3 |
| 7 | DeepSeek R1 Distill Qwen 14B | 3.8 |
| 8 | Llama 3.1 8B Instruct | 1.9 |
Interactive version: theaggregate.ai/benchmark?slug=datagovbench-table-qa · How It Works · Data refreshed daily, snapshot 2026-09-29.