PatternEval - Response-Pattern Failure Rate (Thinking): leaderboard
Metric: Bad-pattern trigger rate (%, lower is better; share of final responses showing at least one of chain-of-thought leakage, repetition, logical contradiction or performative reasoning, labelled by a Seed-2.0-Pro judge with image access; thinking inference; PatternEval: 2,415 failure-enriched image-prompt pairs in three task families (visual perception and grounding, OCR and structured images, multimodal knowledge reasoning); official default decoding with only the reasoning control changed). Source: arxiv.org. Saturation forecast: Estimated already saturated. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 0.78 |
| 2 | Kimi K2.6 (Thinking) | 0.99 |
| 3 | GPT-5.4 (Thinking) | 2.48 |
| 4 | Seed 2.0 Pro (Thinking) | 3.13 |
| 5 | Claude Opus 4.8 (Thinking) | 3.55 |
| 6 | Claude Opus 4.6 (Thinking) | 3.59 |
| 7 | Qwen 3.5 397B A17B (Thinking) | 4.23 |
| 8 | Qwen 3 VL 32B (Thinking) | 4.82 |
| 9 | Qwen 3.5 27B (Thinking) | 5.31 |
| 10 | Qwen 3.7 Max (Thinking) | 5.56 |
| 11 | Qwen 3.5 122B A10B (Thinking) | 5.86 |
| 12 | Qwen 3.6 35B A3B (Thinking) | 5.9 |
| 13 | Qwen 3.6 27B (Thinking) | 5.93 |
| 14 | Kimi K2.5 (Thinking) | 6.14 |
| 15 | Qwen 3.5 35B A3B (Thinking) | 6.37 |
Interactive version: theaggregate.ai/benchmark?slug=patterneval-response-pattern-failure-rate-thinking · How It Works · Data refreshed daily, snapshot 2026-09-29.