VMetaphor-Bench - Metaphoric Efficacy: leaderboard

Metric: Judge rating (1-5) (Qwen3.5-27B rates each generated image from 1 to 5 against the metaphor's structured annotation, 1,500 visual metaphors; this board is the dimension metaphoric efficacy: how clearly the image conveys the intended metaphorical meaning without text; conceptual prompts that state only the metaphorical idea). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1GPT Image 1.54.27
2Nano Banana 2 (Gemini 3.1 Flash Image Preview)4.05
3Seedream 5.0 Lite3.78
4FLUX.2-dev3.48
5Qwen-Image-25123.39
6Stable Diffusion 3.5 Large2.99
7Z-Image-Turbo2.93
8FLUX.1-dev2.9
9Janus-Pro-7B2.9
10BAGEL-7B-MoT2.69

Interactive version: theaggregate.ai/benchmark?slug=vmetaphor-bench-metaphoric-efficacy · How It Works · Data refreshed daily, snapshot 2026-09-26.