InEdit-Bench - Scientific Simulation: leaderboard
Metric: Average (0-100) of the five dimensions (appearance consistency, perceptual quality, semantic consistency, logical coherence, scientific plausibility) on InEdit-Bench's scientific simulation tasks, where the editing model generates the intermediate stages of a multi-step transformation from an initial image to a described final state, scored 0-100 by a GPT-4o-2024-11-20 evaluator; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-Image-1 | 82.61 | |
| 2 | Gemini 2.5 Flash Image (Preview) (Nano Banana) | 79.57 | |
| 3 | Flux-Kontext-pro | 51.3 | |
| 4 | Qwen-Image-Edit | 44.13 | |
| 5 | BAGEL-7B-MoT (Thinking) | 43.48 | |
| 6 | BAGEL-7B-MoT | 39.13 | |
| 7 | OmniGen2 | 34.35 | |
| 8 | Step1X-Edit-v1p1 | 33.91 | |
| 9 | SeedEdit 3.0 | 33.48 | |
| 10 | Emu2 (InEdit-Bench checkpoint unspecified) | 31.3 |
Interactive version: theaggregate.ai/benchmark?slug=inedit-bench-scientific-simulation · How It Works · Data refreshed daily, snapshot 2026-10-11.