VMetaphor-Bench - Metaphor Logic: leaderboard

Metric: Judge rating (1-5) (Qwen3.5-27B rates each generated image from 1 to 5 against the metaphor's structured annotation, 1,500 visual metaphors; this board is the dimension metaphor logic: whether the cross-domain mapping between source and target is conceptually coherent and logically sound; conceptual prompts that state only the metaphorical idea). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1GPT Image 1.54.23
2Nano Banana 2 (Gemini 3.1 Flash Image Preview)4.08
3Seedream 5.0 Lite3.74
4FLUX.2-dev3.46
5Qwen-Image-25123.31
6Z-Image-Turbo3.03
7Janus-Pro-7B3.02
8Stable Diffusion 3.5 Large3.01
9FLUX.1-dev2.87
10BAGEL-7B-MoT2.72

Interactive version: theaggregate.ai/benchmark?slug=vmetaphor-bench-metaphor-logic · How It Works · Data refreshed daily, snapshot 2026-09-26.