SciFlow-Bench - Text Score: leaderboard
Metric: SciFlow-Bench text score (0-1) over 500 scientific framework figures from 2025 arXiv papers: the model draws a diagram from a structured prompt derived from the paper's method section, and an automatic multi-agent parser turns the image back into a graph compared with the figure's canonical graph; coverage (weight 0.3) and faithfulness (weight 0.3) of the recovered components against the structured prompt plus a GPT-4o judge's global alignment (weight 0.4); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Nano Banana Pro | 0.35 | |
| 2 | Gemini 2.5 Flash Image | 0.32 | |
| 3 | Qwen-Image | 0.24 | |
| 4 | PixArt-Sigma | 0 | |
| 5 | SDXL | 0 |
Interactive version: theaggregate.ai/benchmark?slug=sciflow-bench-text-score · How It Works · Data refreshed daily, snapshot 2026-10-11.