MM-CondChain - Natural: leaderboard

Metric: Path F1 (%), the harmonic mean of True-path accuracy (follow every condition to the final answer) and False-path accuracy (stop at the perturbed condition and pick its auxiliary answer), on 398 natural images (SAM and GQA); zero-shot multiple choice with a boxed answer, unparseable outputs wrong; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro55.91#77
2Kimi K2.553.21#139
3Qwen 3 VL 235B A22B Instruct51.47#264
4Qwen 3 VL 235B A22B (Thinking)49.31#228 (Qwen 3 VL 235B A22B)
5GPT-547.51#91
6Gemini 3 Flash47.19#93
7Gemini 2.5 Pro45.7#145
8Qwen 3 VL 8B (Thinking)40.58
9Qwen 3.5 397B A17B38.97#141
10GLM-4.6V38.54#309
11Qwen 3 VL 8B Instruct37.52#401
12Gemini 2.5 Flash36.53#237
13Qwen 3.5 122B A10B34.23#170
14InternVL3-38B32.2#395
15Qwen 3 VL 30B A3B (Thinking)31.03#338 (Qwen 3 VL 30B A3B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mm-condchain-natural · How It Works · Data refreshed daily, snapshot 2026-10-11.