EditReward-Compass - Visual Consistency: leaderboard

Metric: Preference accuracy (%) on the pairs that differ in visual consistency on EditReward-Compass (2,251 human-verified preference pairs of edited images sampled from the same editing models under the same instruction), general MLLMs prompted with the Edit-Compass rubric; thinking disabled unless the row is a thinking run; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)60.02
2Gemini 3 Flash59.81
3Qwen 3.5 27B (Non-reasoning)58.5
4Qwen 3.6 27B56.56
5Qwen 3.5 27B56.37
6Qwen 3.5 35B A3B (Non-reasoning)54.79
7Qwen 3.6 35B A3B53
8Gemma 4 31B51.65
9Qwen 3.5 9B (Non-reasoning)50.75
10Qwen 3.5 35B A3B50.73
11Qwen 3.6 27B (Non-reasoning)49.66
12Qwen 3.5 9B48.6
13GPT-4.148.45
14Qwen 3.6 35B A3B (Non-reasoning)45.58
15Gemma 4 26B A4B39.6

Interactive version: theaggregate.ai/benchmark?slug=editreward-compass-visual-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.