SciFlow-Bench - Image Score: leaderboard
Metric: SciFlow-Bench image score (0-1) over 500 scientific framework figures from 2025 arXiv papers: the model draws a diagram from a structured prompt derived from the paper's method section, and an automatic multi-agent parser turns the image back into a graph compared with the figure's canonical graph; CLIP similarity to the prompt (weight 0.4), a GPT-4o judge's visual-flow consistency (weight 0.4) and inverted LPIPS similarity to the source figure (weight 0.2); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 0.54 | |
| 2 | Gemini 2.5 Flash Image | 0.46 | |
| 3 | Qwen-Image | 0.41 | |
| 4 | PixArt-Sigma | 0.29 | |
| 5 | SDXL | 0.27 |
Interactive version: theaggregate.ai/benchmark?slug=sciflow-bench-image-score · How It Works · Data refreshed daily, snapshot 2026-10-11.