SimoBench v1: leaderboard
Metric: Mean proof score (0–7; 126 synthetic olympiad problems; zero-credit failures included). Source: www.ulam.ai. Saturation forecast: Around January 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (xHigh) | 5.97 |
| 2 | Hy3-preview | 5.9 |
| 3 | Command A+ | 4.44 |
| 4 | DeepSeek V4 Flash | 4.43 |
| 5 | MiniMax-M3 | 3.97 |
Interactive version: theaggregate.ai/benchmark?slug=simobench-v1 · How It Works · Data refreshed daily, snapshot 2026-10-10.