SciImpact: leaderboard

Metric: Pairwise prediction accuracy (%, times 100) on the seven impact dimensions (unweighted mean of the dimension accuracies), over SciImpact contrastive pairs (215,928 pairs over 19 fields, built from citation counts, best paper awards and Nobel prizes, patent citations, media attention, GitHub stars and Hugging Face downloads with impact thresholds): the model sees two same-field artifacts (title and abstract, README, dataset card or model card, truncated to 1,000 words) and names the one with higher future impact in a two-option forced choice parsed by exact match; the order is balanced so chance is 50%; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 11 models tracked.

Top models

#ModelScore
1O4 Mini64.2
2Claude Haiku 4.563.4
3GPT-4.1 Mini62.6
4Qwen 2.5 14B59.2
5Qwen 3 4B58.9
6Qwen 2.5 7B58.9
7Llama 3.1 8B56.4
8Llama 3 8B56.2
9Ministral 3 3B55.8
10nemotron-3-nano-30B-a3B54.3
11Llama 3.2 3B53

Interactive version: theaggregate.ai/benchmark?slug=sciimpact · How It Works · Data refreshed daily, snapshot 2026-10-07.