FED-Bench (Dense Instructions): leaderboard

Metric: FED-Score (x100): product of the alignment (mean of instruction consistency and expression match to the ground-truth image, both judged by Gemini-2.5-Pro), fidelity (mean of ArcFace identity similarity, background error outside the face and judged perceptual quality) and relative expression gain (face LPIPS change over the ground-truth change, a Gaussian penalty centred on 1 that punishes too little and too much change) scores, over FED-Bench's 747 in-the-wild face editing triplets (source image, instruction, ground-truth target) across seven basic expressions, default inference settings; dense instructions describing the facial muscle movements; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 17 models tracked.

Top models

#ModelScoreOverall rank
1FED-Bench Qwen-Image-Edit-Plus (checkpoint unspecified)46.9
2Seedream 4.037.9
3FLUX 2 Pro37.7
4Qwen-Image-Edit33.7
5FLUX-Kontext-Pro32.7
6FLUX-Kontext-Max32
7Qwen-Image-Edit-251131.7
8Step1X-Edit-v1p230.3
9FLUX.1-Kontext-dev24.3
10SeedEdit 3.020.3

Interactive version: theaggregate.ai/benchmark?slug=fed-bench-dense-instructions · How It Works · Data refreshed daily, snapshot 2026-10-11.