QuantEval - Knowledge QA (CoT): leaderboard

Metric: Accuracy (%; chain-of-thought prompting). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.586
2GPT-583.9
3Qwen 3 14B75.9
4Gemini 2.5 Pro66.7
5DeepSeek R1 Distill Qwen 14B66.7
6Qwen 3 8B65.5
7Qwen 3 4B64.4
8DeepSeek R1 Distill Llama 8B50.6
9Qwen 3 30B A3B48.5
10DeepSeek-R1-Distill-Qwen-7B42.5
11DeepSeek R1 Distill Qwen 1.5B37.2

Interactive version: theaggregate.ai/benchmark?slug=quanteval-knowledge-qa-cot · How It Works · Data refreshed daily, snapshot 2026-09-25.