SciUniverse: leaderboard
C5R's benchmark of hands-on scientific work: models direct chemistry, biology and materials experiments at the Facility-0 lab, controlling instruments and instructing human operators, or in its digital twin. Level 1 has 92 tasks in 17 families, from powder pressing, synthesis and liquid-handler control to NMR, XRD and LC-MRM interpretation, reaction optimization under a budget and two simulated weeks of lab management. Pass@1, with each family weighted equally.
Metric: Pass@1 (%). Source: c5r.net. Saturation forecast: Around May 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 (xHigh) | 45.3 |
| 2 | GPT-6 (xHigh) | 32.5 |
| 3 | Claude Opus 5 (xHigh) | 30.5 |
| 4 | Grok 4.6 (xHigh) | 26.2 |
| 5 | Gemini 3.8 Flash (High) | 14.6 |
| 6 | GPT-5.6 Sol (xHigh) | 9.4 |
Interactive version: theaggregate.ai/benchmark?slug=sciuniverse · How It Works · Data refreshed daily, snapshot 2026-09-25.