SciIR-Bench: leaderboard
Metric: Final accuracy (%) over the four tracks on 800 scientific illustration prompts from open-access publications; a gemini-3-pro-preview reviewer answers the atomic binary checks generated for each prompt with visual evidence retrieval; a sample counts only if it passes every check of the track; averaged over intrinsic-reasoning (abstract prompt) and instruction-following (dense scientific reasoning chain prompt) samples; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Nano Banana Pro | 95 |
| 2 | GPT-Image-1 | 62 |
| 3 | Seedream 4.5 | 55 |
| 4 | Qwen-Image-2512 | 35 |
| 5 | FLUX-Kontext-Max | 22 |
| 6 | Show-o2-7B | 21 |
| 7 | HiDream-I1-Full | 13 |
| 8 | FLUX.1-dev | 9 |
| 9 | Stable Diffusion 3.5 Large | 5 |
| 10 | BAGEL-7B-MoT | 2 |
Interactive version: theaggregate.ai/benchmark?slug=sciir-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.