PhysEditBench - Roughness: leaderboard

Metric: Roughness RMSE on the 0-1 scale against the reference roughness map, direct RGB-to-map prompt, OpenRooms-FF (1,125 images); prompted image editors on the paper main-axis split (indoor OpenRooms-FF and InteriorVerse scenes; official access settings of Appendix Table 19); task-specific specialist models are reference rows, not published; lower is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1GPT-Image-20.32
2doubao-seedream-5-0-2601280.36
3GPT-Image-1.50.38
4Qwen-Image-2.00.38

Interactive version: theaggregate.ai/benchmark?slug=physeditbench-roughness · How It Works · Data refreshed daily, snapshot 2026-10-07.