MultiBind - Appearance Binding: leaderboard

Metric: Subject-level success rate (%) in the appearance dimension on MultiBind (multi-subject image generation from per-subject reference images, a background reference and a long entity-indexed prompt, reconstructing a real target photo): share of generated subjects that stay consistent with their own subject in the target photo (Qwen3-VL embeddings, thresholds calibrated to human labels) without being confused with another subject, over the subject slots matched in every model's output; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro95.4
2GPT-Image-1.594.5
3Seedream 4.591.8
4HunyuanImage-3.0-Instruct78
5OmniGen257.5
6Qwen-Image-Edit-251148.4

Interactive version: theaggregate.ai/benchmark?slug=multibind-appearance-binding · How It Works · Data refreshed daily, snapshot 2026-10-11.