MMGist - Diagram and OCR: leaderboard

Metric: Accuracy (%) on the Diagram and OCR dimension (items kept from AI2D and OCRBench); MMGist: 7,262 curated items from 18 vision-language benchmarks after removing items answerable without the image, items nearly every model solves and items with faulty labels; eight samples per item at temperature 1.0, step-by-step answer in a boxed block, 16,384 output tokens, medium effort where configurable; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 27 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Medium)82.3
2Gemini 3.1 Flash Lite (Medium)73.4
3Qwen 3.6 Plus68.7
4Seed 2.0 Pro65.8
5Qwen 3.5 9B65.5
6Qwen 3.6 35B A3B65.4
7Gemma 4 31B63.8
8Seed 2.0 Lite63.2
9Seed 2.0 Mini63.2
10Step3 VL 10B63.1
11Gemma 4 26B A4B61.3
12Claude Sonnet 4.6 (Medium)61.3
13GPT-5 (Medium)58.6
14Qwen 3.5 27B55.4
15GPT-5 Mini (Medium)54.4

Interactive version: theaggregate.ai/benchmark?slug=mmgist-diagram-and-ocr · How It Works · Data refreshed daily, snapshot 2026-09-29.