PhyEditBench - Anti-Physics: leaderboard

Metric: Overall judge score (1-10 scale) on the 35 synthetic anti-physics instances whose edit prompt imposes a counterfactual physical rule (same weighting); GPT-4o judge scores each output 1-10 on consistency, instruction following, physical plausibility and image quality against the instruction, physical explanation, invariants and reference frames; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 12 models tracked.

Top models

#ModelScore
1gemini-2.5-flash-image7.07
2GPT-Image-1.57.04
3Seedream4.06.95
4Qwen-Image-Edit6.16
5BAGEL-7B-MoT (Thinking)6.16
6FLUX.1-Kontext-dev6.07
7UniWorld-V25.79
8Step1X-Edit-v1p15.34
9BAGEL-7B-MoT5.11
10ChronoEdit-14B4.99

Interactive version: theaggregate.ai/benchmark?slug=phyeditbench-anti-physics · How It Works · Data refreshed daily, snapshot 2026-09-29.