SciImpact - Medicine: leaderboard

Metric: Pairwise prediction accuracy (%, times 100) on medicine pairs, over SciImpact contrastive pairs (215,928 pairs over 19 fields, built from citation counts, best paper awards and Nobel prizes, patent citations, media attention, GitHub stars and Hugging Face downloads with impact thresholds): the model sees two same-field artifacts (title and abstract, README, dataset card or model card, truncated to 1,000 words) and names the one with higher future impact in a two-option forced choice parsed by exact match; the order is balanced so chance is 50%; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.

Top models

#ModelScore
1Claude Haiku 4.573
2O4 Mini71
3GPT-4.1 Mini67.3
4Qwen 3 4B60.7
5Qwen 2.5 14B59.7
6Qwen 2.5 7B57.9
7Ministral 3 3B57.7
8Llama 3 8B55.6
9Llama 3.1 8B55.2
10nemotron-3-nano-30B-a3B53.3
11Llama 3.2 3B52.5

Interactive version: theaggregate.ai/benchmark?slug=sciimpact-medicine · How It Works · Data refreshed daily, snapshot 2026-10-07.