OMIBench - Chemistry: leaderboard

Metric: GPTScore (%, binarized LLM-judge equivalence of the final answer with the gold answer, micro-averaged) on the 217 chemistry problems of OMIBench (Olympiad-level multi-image problems from biology, chemistry, mathematics and physics competitions), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 37 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B (Thinking)33.18
2O4 Mini32.41
3Qwen 3 VL 30B A3B (Thinking)32.26
4GPT-529.03
5Gemini 3 Pro (Preview)25.35
6Qwen 3 VL 32B (Thinking)24.88
7GPT-5 Mini24.42
8Gemini 2.5 Pro23.96
9GPT-4o22.58
10Qwen 3 VL 235B A22B Instruct22.58
11GPT-4o Mini21.66
12InternVL3-78B20.74
13InternVL3-14B20.74
14Qwen 3 VL 32B Instruct20.74
15Qwen 2.5 VL 72B Instruct19.82

Interactive version: theaggregate.ai/benchmark?slug=omibench-chemistry · How It Works · Data refreshed daily, snapshot 2026-10-07.