ReLE - Reasoning and Mathematics: leaderboard

Metric: Accuracy (%). Source: nonelinear.com. 174 models tracked.

Top models

#ModelScore
1Claude Opus 592
2Claude Opus 4.8 (Thinking)89.9
3Qwen 3.8 Flash87.7
4Kimi K386.3
5Seed 2.0 Lite85.8
6Qwen 3.5 122B A10B85.5
7Gemini 3.1 Pro (Preview)85.1
8GLM-5.3 Flash84.8
9GPT-5.2 (High)84.8
10Qwen 3.7 Max84.7
11GPT-5.1 (High)84.7
12Gemini 3.5 Flash84.5
13Qwen 3.7 Plus84.5
14GLM-5.384.5
15Gemini 3.8 Flash83.9

Interactive version: theaggregate.ai/benchmark?slug=rele-reasoning-and-mathematics · How It Works · Data refreshed daily, snapshot 2026-09-19.