COHERENCE - Partial Match: leaderboard

Metric: Partial-match score (Kendall's tau between predicted and true order, rescaled to 0-100), over all 6,161 interleaved image-text documents whose images were removed: the model assigns each candidate image to its placeholder; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)90.11
2Claude Sonnet 4.6 (Thinking)88.76
3Qwen 3.5 397B A17B88.37
4Kimi K2.586.74
5GPT-5.4 (High)86.54
6Qwen 3.5 122B A10B86.52
7Qwen 3.5 35B A3B85.06
8Doubao-Seed-2.0-Pro-26021584.76
9Qwen 3 VL 235B A22B81.61
10Qwen 3.5 4B81.16
11GLM-4.6V79.29
12Step3 VL 10B76.75
13Qwen 3 VL 8B Instruct75.86
14GPT-5.274.84
15Qwen 3 VL 4B Instruct72.9

Interactive version: theaggregate.ai/benchmark?slug=coherence-partial-match · How It Works · Data refreshed daily, snapshot 2026-10-07.