VMetaphor-Bench - Meaning: leaderboard

Metric: MCQ accuracy (%; one question per image on the overall metaphorical message; the generated image is judged by Qwen3.5-27B on multiple-choice questions (four options including cannot be determined) built from each metaphor's structured annotation; 1,500 visual metaphors curated from real creative imagery; conceptual prompts that state only the metaphorical idea). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1Nano Banana 2 (Gemini 3.1 Flash Image Preview)98.5
2GPT Image 1.597.6
3Seedream 5.0 Lite96
4FLUX.2-dev92.1
5Qwen-Image-251290.5
6Z-Image-Turbo89.5
7Janus-Pro-7B89.1
8Stable Diffusion 3.5 Large85.7
9FLUX.1-dev85.6
10BAGEL-7B-MoT84.2

Interactive version: theaggregate.ai/benchmark?slug=vmetaphor-bench-meaning · How It Works · Data refreshed daily, snapshot 2026-09-26.