SciImpact - Code: leaderboard

Metric: Pairwise prediction accuracy (%, times 100) on code repository pairs (GitHub stars, from READMEs), over SciImpact contrastive pairs (215,928 pairs over 19 fields, built from citation counts, best paper awards and Nobel prizes, patent citations, media attention, GitHub stars and Hugging Face downloads with impact thresholds): the model sees two same-field artifacts (title and abstract, README, dataset card or model card, truncated to 1,000 words) and names the one with higher future impact in a two-option forced choice parsed by exact match; the order is balanced so chance is 50%; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 11 models tracked.

Top models

#ModelScore
1O4 Mini65.8
2Claude Haiku 4.562.6
3GPT-4.1 Mini60.3
4Qwen 2.5 14B57.7
5Qwen 3 4B57.3
6Qwen 2.5 7B56.3
7Llama 3 8B54.7
8Ministral 3 3B54.2
9nemotron-3-nano-30B-a3B52.8
10Llama 3.1 8B52.5
11Llama 3.2 3B51.3

Interactive version: theaggregate.ai/benchmark?slug=sciimpact-code · How It Works · Data refreshed daily, snapshot 2026-10-07.