ReactBench (Reaction Diagrams): leaderboard

Metric: Answer accuracy (%) on the four task dimensions (unweighted mean of localization, extraction, tracing and reasoning accuracy), over the 1,618 expert-annotated question-answer pairs on more than 1,300 real chemical reaction diagrams from chemistry journals and patents of ReactBench (single-line, multi-line, tree and graph topologies), multiple-choice or short open answers matched exactly against the key after template extraction, no LLM judge; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 24 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct73.91
2Gemini 3.1 Pro (Preview)73.88
3Qwen 3 VL 32B (Thinking)70.89
4GPT-5.569.59
5Claude 3.5 Sonnet68.6
6Claude Sonnet 4.667
7GPT-4o63.83
8Qwen 3 VL 8B Instruct63.02
9Gemini 1.5 Pro62.17
10Qwen 3 VL 8B (Thinking)60.24
11Qwen 2.5 VL 7B57.32
12InternVL2.5-78B53.74

Interactive version: theaggregate.ai/benchmark?slug=reactbench-reaction-diagrams · How It Works · Data refreshed daily, snapshot 2026-10-07.