Equational Theories - Hard (Default reasoning): leaderboard

Metric: Strict F1 (%). Source: huggingface.co. Saturation forecast: Around January 2027. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)87.84
2Grok 4.1 Fast70.45
3Qwen 3.5 397B A17B65.51
4Kimi K2.565
5Qwen 3.5 122B A10B60.82
6Seed 2.0 Lite60.14
7Qwen 3.5 27B60
8GPT-5 Mini57.93
9GPT-OSS-120B56.73
10GLM-554.22
11Claude Sonnet 4.649.28
12Step 3.5 Flash46.43
13MiniMax-M2.546.27
14Llama 3.1 8B Instruct39.7
15Grok Code Fast 138.96

Interactive version: theaggregate.ai/benchmark?slug=equational-theories-hard-default-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.