Equational Theories - Normal (Default reasoning): leaderboard

Metric: Strict F1 (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1GPT-5.497.98
2Gemini 3.1 Pro (Preview)97.81
3Qwen 3.5 397B A17B97.37
4Grok 4.1 Fast96.87
5Qwen 3.5 122B A10B96.79
6Qwen 3.5 27B94.97
7Gemini 3 Flash (Preview)94.1
8Kimi K2.593.59
9GPT-5 Mini93.44
10GPT-OSS-120B92.59
11GLM-591.73
12Seed 2.0 Lite90.52
13Step 3.5 Flash83.54
14Claude Sonnet 4.683
15MiniMax-M2.579.86

Interactive version: theaggregate.ai/benchmark?slug=equational-theories-normal-default-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.