SciEvalKit — leaderboard

Scientific capability ranking from OpenCompass and SciEvalKit, measuring frontier models on science-heavy reasoning and knowledge tasks.

Metric: Scientific Capability Score. Source: github.com. Status: years away from saturation. 10 models tracked.

Top models

#ModelScore
1Gemini 3 Pro48.74
2Claude Opus 4.545.78
3Qwen 3 Max45.33
4GPT-545.12
5Claude Sonnet 4.542.12
6Seed 1.842.05
7Qwen 3 VL 235B A22B41.43
8DeepSeek V3.241.35
9O341.08
10Gemini 2.5 Pro40.77

Interactive version: theaggregate.ai/benchmark?slug=scievalkit · How the rankings work · Data refreshed daily, snapshot 2026-07-22.