MMTU — leaderboard

Massive Multi-Task Table Understanding benchmark with real-world table reasoning tasks including joins, transformations, cleaning, and question answering.

Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 26 models tracked.

Top models

#ModelScore
1GPT-569.58
2O369.13
3GPT-5 Mini66.74
4Gemini 2.5 Pro66.45
5Grok 3 Mini64.55
6Gemini 2.5 Flash62.55
7DeepSeek R1 052857.99
8GPT-5 Chat57.67
9GPT-OSS-120B54.32
10Qwen 3 235B A22B (Thinking)52.87
11Qwen 3 235B A22B 2507 Instruct52.42
12GPT-4o (2024-11-20)50.73
13Qwen 3 32B50.6
14Llama 4 Maverick Instruct FP849
15GPT-OSS-20B47.8

Interactive version: theaggregate.ai/benchmark?slug=mmtu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.