RS-RIE-Bench: leaderboard

Metric: Strict joint-satisfaction accuracy (%; share of the 486 reasoning-guided remote sensing image editing tasks whose edited image scores 5 out of 5 on target-region plausibility, non-target region preservation and image-quality consistency, judged by gpt-5.1 with a fixed rubric; higher is better). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1gpt-image-224.28
2doubao-seedream-5-0-26012820.37
3gemini-3.1-flash-image-preview12.76
4FLUX.2-dev6.79
5wan2.7-image-pro5.97
6grok-imagine-image3.09
7Step1X-Edit0.82
8Qwen-Image-Edit-25090.62

Interactive version: theaggregate.ai/benchmark?slug=rs-rie-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.