FED-Bench (Simple Instructions): leaderboard
Metric: FED-Score (x100): product of the alignment (mean of instruction consistency and expression match to the ground-truth image, both judged by Gemini-2.5-Pro), fidelity (mean of ArcFace identity similarity, background error outside the face and judged perceptual quality) and relative expression gain (face LPIPS change over the ground-truth change, a Gaussian penalty centred on 1 that punishes too little and too much change) scores, over FED-Bench's 747 in-the-wild face editing triplets (source image, instruction, ground-truth target) across seven basic expressions, default inference settings; simple instructions such as change the expression from neutral to happy; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 17 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | FED-Bench Qwen-Image-Edit-Plus (checkpoint unspecified) | 49.2 | |
| 2 | Seedream 4.0 | 41.3 | |
| 3 | FLUX 2 Pro | 40 | |
| 4 | Qwen-Image-Edit-2511 | 36.1 | |
| 5 | Qwen-Image-Edit | 34.3 | |
| 6 | Step1X-Edit-v1p2 | 33.3 | |
| 7 | FLUX-Kontext-Max | 25.9 | |
| 8 | SeedEdit 3.0 | 23.9 | |
| 9 | FLUX-Kontext-Pro | 22.7 | |
| 10 | BAGEL-7B-MoT | 16.3 |
Interactive version: theaggregate.ai/benchmark?slug=fed-bench-simple-instructions · How It Works · Data refreshed daily, snapshot 2026-10-11.