GeoR-Bench - Consistency: leaderboard

Metric: Consistency score (0-100): whether the output preserves the visual relationship to the input image that the task requires (mean of the six category scores), 440 reasoning-informed geoscience image-editing samples in six categories (geomorphology, hydrology, atmosphere and ocean, cryosphere, GIS and spatial geometry, crustal science), one generated output per sample, Gemini 3 Flash judge with task-specific rubrics; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 21 models tracked.

Top models

#ModelScore
1Nano Banana Pro96.3
2GPT-Image-295.1
3Nano Banana 294.6
4GPT-Image-1.594.4
5Seedream 5.090.3
6Nano Banana88.1
7FLUX 2 Max71.6
8Seedream 4.571.3
9Qwen-Image-Edit-251162.6
10FLUX.2-dev62.2

Interactive version: theaggregate.ai/benchmark?slug=geor-bench-consistency · How It Works · Data refreshed daily, snapshot 2026-10-07.