TangPoetryBench: leaderboard

Metric: Human quality score (0-100): per-image mean over the applicable of ten rated dimensions, from visual quality and poem correspondence to imagery, emotion and overall impression, averaged across raters and over 320 poem illustrations. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1Nano Banana Pro91.1
2GPT Image 189.2
3Seedream 4.588
4Midjourney V772.5

Interactive version: theaggregate.ai/benchmark?slug=tangpoetrybench · How It Works · Data refreshed daily, snapshot 2026-09-26.