SciFigQual-Bench: leaderboard

Metric: Mean absolute error of the overall figure-quality score against the mean expert score (1-10 rating scale, lower is better; 1,200 figures from 254 ACL, EMNLP, ICML and NeurIPS papers, each rated with its caption and citing paragraphs; one direct VLM call per figure). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.80.44
2GPT-5.6 Sol0.45
3Gemini 3.5 Flash0.46
4Nova Pro0.47
5Claude Sonnet 50.48
6Llama 4 Maverick0.49
7Pixtral Large0.5
8InternVL3-78B0.54
9Seed 2.0 Pro0.61
10GLM-4.6V0.66

Interactive version: theaggregate.ai/benchmark?slug=scifigqual-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.