MultiBind - Appearance Binding: leaderboard
Metric: Subject-level success rate (%) in the appearance dimension on MultiBind (multi-subject image generation from per-subject reference images, a background reference and a long entity-indexed prompt, reconstructing a real target photo): share of generated subjects that stay consistent with their own subject in the target photo (Qwen3-VL embeddings, thresholds calibrated to human labels) without being confused with another subject, over the subject slots matched in every model's output; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 95.4 | |
| 2 | GPT-Image-1.5 | 94.5 | |
| 3 | Seedream 4.5 | 91.8 | |
| 4 | HunyuanImage-3.0-Instruct | 78 | |
| 5 | OmniGen2 | 57.5 | |
| 6 | Qwen-Image-Edit-2511 | 48.4 |
Interactive version: theaggregate.ai/benchmark?slug=multibind-appearance-binding · How It Works · Data refreshed daily, snapshot 2026-10-11.