QuantEval - Quantitative Reasoning (CoT): leaderboard

Metric: Accuracy (%; chain-of-thought prompting). Source: arxiv.org. Saturation forecast: Around 2029. 13 models tracked.

Top models

#ModelScore
1GPT-555
2Claude Sonnet 4.543.1
3Qwen 3 30B A3B40.5
4Gemini 2.5 Pro38
5Qwen 3 14B35.4
6Qwen 3 8B33.9
7DeepSeek R1 Distill Qwen 1.5B30
8DeepSeek R1 Distill Qwen 14B29.2
9Qwen 3 4B27.7
10DeepSeek-R1-Distill-Qwen-7B21.5
11DeepSeek R1 Distill Llama 8B10.8

Interactive version: theaggregate.ai/benchmark?slug=quanteval-quantitative-reasoning-cot · How It Works · Data refreshed daily, snapshot 2026-09-25.