VMetaphor-Bench (Descriptive Prompts): leaderboard

Metric: MCQ accuracy (%; pooled over all 9,594 fidelity questions; the generated image is judged by Qwen3.5-27B on multiple-choice questions (four options including cannot be determined) built from each metaphor's structured annotation; 1,500 visual metaphors curated from real creative imagery; descriptive prompts that spell out the visual realization). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1Nano Banana 2 (Gemini 3.1 Flash Image Preview)89.2
2GPT Image 1.588.5
3FLUX.2-dev87.3
4Qwen-Image-251287
5Seedream 5.0 Lite86.6
6Z-Image-Turbo85.9
7Janus-Pro-7B81
8FLUX.1-dev80.7
9BAGEL-7B-MoT79.1
10Stable Diffusion 3.5 Large78.8

Interactive version: theaggregate.ai/benchmark?slug=vmetaphor-bench-descriptive-prompts · How It Works · Data refreshed daily, snapshot 2026-09-26.