PatternEval (Thinking): leaderboard

Metric: Accuracy (%; correctness of the final user-visible answer under the source task's criteria, judged by Qwen3-Max; thinking inference; PatternEval: 2,415 failure-enriched image-prompt pairs in three task families (visual perception and grounding, OCR and structured images, multimodal knowledge reasoning); official default decoding with only the reasoning control changed). Source: arxiv.org. Saturation forecast: Around January 2027. 25 models tracked.

Top models

#ModelScore
1GPT-5.574.9
2Kimi K2.6 (Thinking)73.19
3Seed 2.0 Pro (Thinking)72.98
4Qwen 3.7 Max (Thinking)72.3
5GPT-5.4 (Thinking)69.54
6Qwen 3.5 397B A17B (Thinking)68.93
7Claude Opus 4.8 (Thinking)67.7
8Claude Opus 4.6 (Thinking)67.63
9Kimi K2.5 (Thinking)67.06
10Qwen 3.5 122B A10B (Thinking)66.3
11Qwen 3.5 35B A3B (Thinking)65.07
12Qwen 3.5 27B (Thinking)64.05
13Qwen 3.6 35B A3B (Thinking)63.9
14Qwen 3.6 27B (Thinking)63.72
15Qwen 3.5 9B (Thinking)59.2

Interactive version: theaggregate.ai/benchmark?slug=patterneval-thinking · How It Works · Data refreshed daily, snapshot 2026-09-29.