SciFlow-Bench - Text Score: leaderboard

Metric: SciFlow-Bench text score (0-1) over 500 scientific framework figures from 2025 arXiv papers: the model draws a diagram from a structured prompt derived from the paper's method section, and an automatic multi-agent parser turns the image back into a graph compared with the figure's canonical graph; coverage (weight 0.3) and faithfulness (weight 0.3) of the recovered components against the structured prompt plus a GPT-4o judge's global alignment (weight 0.4); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScoreOverall rank
1Nano Banana Pro0.35
2Gemini 2.5 Flash Image0.32
3Qwen-Image0.24
4PixArt-Sigma0
5SDXL0

Interactive version: theaggregate.ai/benchmark?slug=sciflow-bench-text-score · How It Works · Data refreshed daily, snapshot 2026-10-11.