OMIBench - Physics: leaderboard

Metric: GPTScore (%, binarized LLM-judge equivalence of the final answer with the gold answer, micro-averaged) on the 424 physics problems of OMIBench (Olympiad-level multi-image problems from biology, chemistry, mathematics and physics competitions), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 37 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B (Thinking)43.87
2GPT-5 Mini43.63
3Qwen 3 VL 32B (Thinking)41.98
4GPT-540.8
5Gemini 3 Pro (Preview)38.92
6Qwen 3 VL 30B A3B (Thinking)37.5
7O4 Mini35.38
8Gemini 2.5 Pro31.84
9Qwen 3 VL 8B (Thinking)30.66
10Qwen 3 VL 32B Instruct25
11Qwen 3 VL 235B A22B Instruct23.58
12Gemini 2.5 Flash23.35
13Qwen 3 VL 30B A3B Instruct20.99
14Qwen 3 VL 4B (Thinking)20.75
15Qwen 2.5 VL 32B Instruct19.81

Interactive version: theaggregate.ai/benchmark?slug=omibench-physics · How It Works · Data refreshed daily, snapshot 2026-10-07.