Qwen-Image-Bench - Human Expert Rating - Real-World Fidelity: leaderboard

Metric: Real-world fidelity (knowledge-consistent depiction of real entities, text and physical facts) pillar professional annotator rating on a 1 to 10 scale over 1,000 prompts per pillar, averaged; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1GPT Image 28.8
2Nano Banana 28.64
3Nano Banana Pro8.45
4GPT Image 1.57.29
5Qwen Image 2.0 Pro7.27
6Seedream 5.06.22
7Qwen Image 25125.23
8Seedream 4.55.17
9GPT Image 14.98
10Seedream 4.04.67

Interactive version: theaggregate.ai/benchmark?slug=qwen-image-bench-human-expert-rating-real-world-fidelity · How It Works · Data refreshed daily, snapshot 2026-10-07.