SciIR-Bench: leaderboard

Metric: Final accuracy (%) over the four tracks on 800 scientific illustration prompts from open-access publications; a gemini-3-pro-preview reviewer answers the atomic binary checks generated for each prompt with visual evidence retrieval; a sample counts only if it passes every check of the track; averaged over intrinsic-reasoning (abstract prompt) and instruction-following (dense scientific reasoning chain prompt) samples; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.

Top models

#ModelScore
1Nano Banana Pro95
2GPT-Image-162
3Seedream 4.555
4Qwen-Image-251235
5FLUX-Kontext-Max22
6Show-o2-7B21
7HiDream-I1-Full13
8FLUX.1-dev9
9Stable Diffusion 3.5 Large5
10BAGEL-7B-MoT2

Interactive version: theaggregate.ai/benchmark?slug=sciir-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.