TableVista - Multi-Table: leaderboard

Metric: Accuracy (%) on the 700 multi-table questions of the 3,000 TableVista questions, Web rendering, direct-output prompt with thinking disabled (reasoning effort none for the GPT models), exact match with a GPT-5-mini semantic-equivalence check on exact-match failures; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 29 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)61.3
2Qwen 2.5 VL 72B Instruct53.1
3Llama 4 Maverick Instruct FP852.4
4Gemma 4 31B (IT)52.3
5Qwen 2.5 VL 32B Instruct49.1
6Qwen 3.5 27B (Non-reasoning)48.6
7Llama 4 Scout Instruct48.1
8Qwen 3.5 122B A10B (Non-reasoning)46.3
9Qwen 2.5 VL 7B Instruct42.4
10Gemma 4 26B A4B (IT)42
11Qwen 3 VL 30B A3B Instruct41
12Qwen 3.6 35B A3B (Non-reasoning)41
13GPT-5.4 Mini (Non-reasoning)40
14Qwen 3 VL 8B Instruct39.9
15Gemma 3 27B (IT)39.6

Interactive version: theaggregate.ai/benchmark?slug=tablevista-multi-table · How It Works · Data refreshed daily, snapshot 2026-10-07.