Equational Theories - Hard2 (Low-or-none reasoning): leaderboard

Metric: Strict F1 (%). Source: huggingface.co. Saturation forecast: Not forecast. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)91.19
2Qwen 3.5 397B A17B71.52
3Qwen 3.5 122B A10B70.13
4Qwen 3.5 27B69.17
5GPT-OSS-120B64.88
6GPT-5 Mini64.38
7Claude Opus 4.659.71
8Llama 3.1 8B Instruct59.55
9Seed 2.0 Lite58.17
10Claude Sonnet 4.654.82
11Grok Code Fast 150.23
12GPT-5 Nano35.35
13Gemini 3 Flash (Preview)30.53
14MiniMax-M2.524.63
15Claude Haiku 4.518.68

Interactive version: theaggregate.ai/benchmark?slug=equational-theories-hard2-low-or-none-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.