VMetaphor-Bench - Dimension Score (Descriptive Prompts): leaderboard

Metric: Judge rating (1-5) (Qwen3.5-27B rates each generated image from 1 to 5 on metaphoric efficacy, metaphor logic and perceptual harmony against the structured annotation, and the board takes the mean of the three; 1,500 visual metaphors; descriptive prompts that spell out the visual realization). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1GPT Image 1.54.24
2Nano Banana 2 (Gemini 3.1 Flash Image Preview)4.18
3Seedream 5.0 Lite4.15
4FLUX.2-dev4.09
5Qwen-Image-25124.03
6Z-Image-Turbo3.9
7FLUX.1-dev3.66
8Stable Diffusion 3.5 Large3.54
9Janus-Pro-7B3.48
10BAGEL-7B-MoT3.47

Interactive version: theaggregate.ai/benchmark?slug=vmetaphor-bench-dimension-score-descriptive-prompts · How It Works · Data refreshed daily, snapshot 2026-09-26.