Equational Theories - Normal (Low-or-none reasoning): leaderboard

Metric: Strict F1 (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1Qwen 3.5 122B A10B97.53
2GPT-5.497.19
3Gemini 3.1 Pro (Preview)96.23
4Qwen 3.5 397B A17B91.03
5Claude Opus 4.684.75
6Gemini 3 Flash (Preview)84.23
7Claude Sonnet 4.682.8
8GPT-5 Mini81.07
9Seed 2.0 Lite77.4
10Kimi K2.574.03
11GPT-OSS-120B71.74
12Qwen 3.5 27B68.22
13GLM-561.94
14MiniMax-M2.560.34
15GPT-5 Nano51.98

Interactive version: theaggregate.ai/benchmark?slug=equational-theories-normal-low-or-none-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.