SGI-Bench — leaderboard

Scientific General Intelligence benchmark testing LLMs across deep research, idea generation, dry/wet experiments, and experimental reasoning. Evaluates scientific discovery capabilities.

Metric: SGI-Score. Source: github.com. Status: saturation imminent. 19 models tracked.

Top models

#ModelScore
1Gemini 3 Pro33.83
2Claude Sonnet 4.532.16
3Qwen 3 Max31.97
4GPT-4.131.45
5GPT-5.2 Pro31.09
6GPT-530.84
7O330.68
8Claude Opus 4.130.42
9O4 Mini30.14
10GPT-5.129.31
11Grok 428.68
12Qwen 3 VL 235B A22B28.32
13Gemini 2.5 Pro28.17
14Intern-S128.1
15GPT-4o26.87

Interactive version: theaggregate.ai/benchmark?slug=sgi-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.