MultiBind - Expression Binding: leaderboard

Metric: Subject-level success rate (%) in the facial expression dimension on MultiBind (multi-subject image generation from per-subject reference images, a background reference and a long entity-indexed prompt, reconstructing a real target photo): share of generated subjects that stay consistent with their own subject in the target photo (Qwen3-VL embeddings, thresholds calibrated to human labels) without being confused with another subject, over the subject slots matched in every model's output; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro95.3
2Seedream 4.593.9
3GPT-Image-1.593.4
4HunyuanImage-3.0-Instruct87.8
5Qwen-Image-Edit-251168.1
6OmniGen266.9

Interactive version: theaggregate.ai/benchmark?slug=multibind-expression-binding · How It Works · Data refreshed daily, snapshot 2026-10-11.