PatternEval - Response-Pattern Failure Rate (Thinking): leaderboard

Metric: Bad-pattern trigger rate (%, lower is better; share of final responses showing at least one of chain-of-thought leakage, repetition, logical contradiction or performative reasoning, labelled by a Seed-2.0-Pro judge with image access; thinking inference; PatternEval: 2,415 failure-enriched image-prompt pairs in three task families (visual perception and grounding, OCR and structured images, multimodal knowledge reasoning); official default decoding with only the reasoning control changed). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1GPT-5.50.78
2Kimi K2.6 (Thinking)0.99
3GPT-5.4 (Thinking)2.48
4Seed 2.0 Pro (Thinking)3.13
5Claude Opus 4.8 (Thinking)3.55
6Claude Opus 4.6 (Thinking)3.59
7Qwen 3.5 397B A17B (Thinking)4.23
8Qwen 3 VL 32B (Thinking)4.82
9Qwen 3.5 27B (Thinking)5.31
10Qwen 3.7 Max (Thinking)5.56
11Qwen 3.5 122B A10B (Thinking)5.86
12Qwen 3.6 35B A3B (Thinking)5.9
13Qwen 3.6 27B (Thinking)5.93
14Kimi K2.5 (Thinking)6.14
15Qwen 3.5 35B A3B (Thinking)6.37

Interactive version: theaggregate.ai/benchmark?slug=patterneval-response-pattern-failure-rate-thinking · How It Works · Data refreshed daily, snapshot 2026-09-29.