EditReward-Compass: leaderboard

Metric: Average preference accuracy (%) over all pairs on EditReward-Compass (2,251 human-verified preference pairs of edited images sampled from the same editing models under the same instruction), general MLLMs prompted with the Edit-Compass rubric; thinking disabled unless the row is a thinking run; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)74.33
2Gemini 3 Flash72.68
3Qwen 3.6 27B71.83
4Qwen 3.5 35B A3B70.89
5Qwen 3.6 35B A3B70.51
6Qwen 3.5 27B69.98
7Gemma 4 31B67.09
8Qwen 3.5 27B (Non-reasoning)66.93
9Qwen 3.5 9B66.81
10GPT-4.166.11
11Qwen 3.6 27B (Non-reasoning)63.28
12Qwen 3.5 35B A3B (Non-reasoning)63.18
13Qwen 3.5 9B (Non-reasoning)60.16
14Qwen 3.6 35B A3B (Non-reasoning)59.95
15Gemma 4 26B A4B59.6

Interactive version: theaggregate.ai/benchmark?slug=editreward-compass · How It Works · Data refreshed daily, snapshot 2026-10-07.