RS-RIE-Bench: leaderboard
Metric: Strict joint-satisfaction accuracy (%; share of the 486 reasoning-guided remote sensing image editing tasks whose edited image scores 5 out of 5 on target-region plausibility, non-target region preservation and image-quality consistency, judged by gpt-5.1 with a fixed rubric; higher is better). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-image-2 | 24.28 |
| 2 | doubao-seedream-5-0-260128 | 20.37 |
| 3 | gemini-3.1-flash-image-preview | 12.76 |
| 4 | FLUX.2-dev | 6.79 |
| 5 | wan2.7-image-pro | 5.97 |
| 6 | grok-imagine-image | 3.09 |
| 7 | Step1X-Edit | 0.82 |
| 8 | Qwen-Image-Edit-2509 | 0.62 |
Interactive version: theaggregate.ai/benchmark?slug=rs-rie-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.