InEdit-Bench - Scientific Simulation: leaderboard

Metric: Average (0-100) of the five dimensions (appearance consistency, perceptual quality, semantic consistency, logical coherence, scientific plausibility) on InEdit-Bench's scientific simulation tasks, where the editing model generates the intermediate stages of a multi-step transformation from an initial image to a described final state, scored 0-100 by a GPT-4o-2024-11-20 evaluator; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-Image-182.61
2Gemini 2.5 Flash Image (Preview) (Nano Banana)79.57
3Flux-Kontext-pro51.3
4Qwen-Image-Edit44.13
5BAGEL-7B-MoT (Thinking)43.48
6BAGEL-7B-MoT39.13
7OmniGen234.35
8Step1X-Edit-v1p133.91
9SeedEdit 3.033.48
10Emu2 (InEdit-Bench checkpoint unspecified)31.3

Interactive version: theaggregate.ai/benchmark?slug=inedit-bench-scientific-simulation · How It Works · Data refreshed daily, snapshot 2026-10-11.