COHERENCE - Cooking: leaderboard

Metric: Exact-match accuracy (%) on the 1,326 cooking recipe documents; 6,161 interleaved image-text documents whose images were removed: the model assigns each candidate image to its placeholder; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)79.19
2Gemini 3.1 Pro (Preview)73.45
3Claude Sonnet 4.6 (Thinking)73.15
4Qwen 3.5 397B A17B71.27
5Doubao-Seed-2.0-Pro-26021563.27
6Qwen 3.5 122B A10B61.54
7Kimi K2.558.97
8GPT-5.255.28
9Qwen 3.5 35B A3B54.15
10Qwen 3 VL 235B A22B46.83
11Qwen 3.5 4B43.14
12Step3 VL 10B42.84
13GLM-4.6V40.27
14Qwen 3 VL 8B Instruct31.3
15Qwen 3 VL 4B Instruct26.24

Interactive version: theaggregate.ai/benchmark?slug=coherence-cooking · How It Works · Data refreshed daily, snapshot 2026-10-07.