WildTableBench - Color: leaderboard

Metric: Accuracy (%) on the Color question category (C5, question-weighted over its subtypes), 928 human-verified questions over 402 real-world table images, free-form answers judged by GPT-5.2 against the reference answer, high reasoning effort where configurable; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 21 models tracked.

Top models

#ModelScore
1Gemini 3 Pro55.8
2Gemini 3 Flash37.5
3Kimi K2.5 (Thinking)37.5
4Seed 2.0 Pro (High)32.5
5GPT-5.2 (High)29.2
6Qwen 3 VL 32B (Thinking)28.3
7Qwen 3 VL 235B A22B (Thinking)26.7
8GLM-4.6V25.8
9Qwen 3 VL 235B A22B Instruct21.7
10Claude Opus 4.6 (High)20.8
11Qwen 3 VL 32B Instruct18.3
12GPT-5 Mini (High)18.3
13Claude Sonnet 4.6 (High)17.5
14O3 (High)15
15Qwen 3 VL 8B (Thinking)11.7

Interactive version: theaggregate.ai/benchmark?slug=wildtablebench-color · How It Works · Data refreshed daily, snapshot 2026-10-07.