ChemPro (MCQ) - Challenging: leaderboard

Metric: Multiple-choice accuracy (%) on the Challenging section (advanced high-school NCERT questions) of ChemPro; pass@1, mean of five runs, fixed system and wraparound prompts for every model; printed as fractions and shown times 100; Table 2 lists only the three best open models of each size class (7B, 10B, 14B, 32B, 70B) and OpenAI o1-mini, o3-mini and o1; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScoreOverall rank
1O185#177
2O3 Mini84#266
3O1 Mini83#313
4Rombos-LLM-V2.5-Qwen-32B81#353
5Qwen2.5-32B-Instruct-abliterated-v280#381
6Falcon3-10B-Instruct77#1142
7Lamarckvergence-14B77#533
8Falcon3-7B-Instruct70#1284
9HomerCreativeAnvita-Mix-Qw7B67#917

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=chempro-mcq-challenging · How It Works · Data refreshed daily, snapshot 2026-10-11.