PhyEditBench - Physical Plausibility: leaderboard

Metric: Physical plausibility judge score (1-10 scale) on the 238 real-world physical-process instances, averaged over the five run types; GPT-4o judge scores each output 1-10 on consistency, instruction following, physical plausibility and image quality against the instruction, physical explanation, invariants and reference frames; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.

Top models

#ModelScore
1ChronoEdit-14B8.42
2GPT-Image-1.58.3
3Seedream4.08.22
4UniWorld-V26.73
5gemini-2.5-flash-image6.64
6Qwen-Image-Edit6.37
7Step1X-Edit-v1p16.19
8BAGEL-7B-MoT (Thinking)5.96
9BAGEL-7B-MoT5.69
10InstructPix2Pix5.58

Interactive version: theaggregate.ai/benchmark?slug=phyeditbench-physical-plausibility · How It Works · Data refreshed daily, snapshot 2026-09-29.