MM-CondChain - GUI: leaderboard

Metric: Path F1 (%), the harmonic mean of True-path accuracy (follow every condition to the final answer) and False-path accuracy (stop at the perturbed condition and pick its auxiliary answer), on 377 AITZ GUI trajectories (3,421 screenshots); zero-shot multiple choice with a boxed answer, unparseable outputs wrong; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 27 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3.5 397B A17B40.19#141
2GPT-538.06#91
3Gemini 3 Pro38.05#77
4Gemini 3 Flash35.78#93
5Qwen 3.5 122B A10B34.17#170
6Kimi K2.533.72#139
7Qwen 3 VL 30B A3B (Thinking)32.93#338 (Qwen 3 VL 30B A3B)
8Qwen 3 VL 8B (Thinking)31.83
9Qwen 3 VL 235B A22B (Thinking)31.23#228 (Qwen 3 VL 235B A22B)
10GLM-4.6V27.11#309
11Qwen 3 VL 235B A22B Instruct27.04#264
12Qwen 3.5 35B A3B24.02#250
13Qwen 3 VL 8B Instruct20.65#401
14InternVL3-38B20.46#395
15GPT-4o (2024-11-20)20.46#369

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mm-condchain-gui · How It Works · Data refreshed daily, snapshot 2026-10-11.