SciPredict — leaderboard

SciPredict benchmarks LLMs on forecasting the outcomes of real scientific experiments across biology, chemistry, and physics.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.523.05
2Claude Sonnet 4.522.55
3Gemini 3 Flash (Preview)22.22
4Claude Opus 4.122.22
5GPT-5.220.58
6O3 Mini19.84
7Llama 3.3 70B Instruct18.19
8Gemini 2.5 Pro17.04
9Qwen 3 235B A22B16.63
10Llama 3.1 8B14.65

Interactive version: theaggregate.ai/benchmark?slug=scipredict · How the rankings work · Data refreshed daily, snapshot 2026-07-22.