ChemPro (MCQ) - Medium: leaderboard

Metric: Multiple-choice accuracy (%) on the Medium section (intermediate textbook questions) of ChemPro; pass@1, mean of five runs, fixed system and wraparound prompts for every model; printed as fractions and shown times 100; Table 2 lists only the three best open models of each size class (7B, 10B, 14B, 32B, 70B) and OpenAI o1-mini, o3-mini and o1; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScoreOverall rank
1O199#177
2O3 Mini99#266
3Rombos-LLM-V2.5-Qwen-32B99#353
4O1 Mini98#313
5Qwen2.5-32B-Instruct-abliterated-v298#381
6Lamarckvergence-14B98#533
7Falcon3-10B-Instruct97#1142
8Falcon3-7B-Instruct94#1284
9HomerCreativeAnvita-Mix-Qw7B93#917

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=chempro-mcq-medium · How It Works · Data refreshed daily, snapshot 2026-10-11.