EditReward-Compass - Instruction Awareness: leaderboard

Metric: Preference accuracy (%) on the pairs that differ in instruction awareness on EditReward-Compass (2,251 human-verified preference pairs of edited images sampled from the same editing models under the same instruction), general MLLMs prompted with the Edit-Compass rubric; thinking disabled unless the row is a thinking run; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)83.24
2Qwen 3.5 35B A3B80.74
3Gemini 3 Flash80.42
4Qwen 3.6 27B79.61
5Qwen 3.6 35B A3B79.21
6Qwen 3.5 27B76.74
7Qwen 3.5 9B76.15
8Gemma 4 31B75.27
9GPT-4.174.71
10Qwen 3.5 27B (Non-reasoning)73.22
11Qwen 3.5 35B A3B (Non-reasoning)72.79
12Qwen 3.6 27B (Non-reasoning)71.47
13Gemma 4 26B A4B69.47
14Qwen 3.6 35B A3B (Non-reasoning)68.24
15Qwen 3.5 9B (Non-reasoning)66.82

Interactive version: theaggregate.ai/benchmark?slug=editreward-compass-instruction-awareness · How It Works · Data refreshed daily, snapshot 2026-10-07.