QuantEval - Knowledge QA (No CoT): leaderboard

Metric: Accuracy (%; direct answer, no chain-of-thought). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.590.8
2GPT-582.8
3Qwen 3 14B78.2
4Qwen 3 4B69
5Gemini 2.5 Pro66.7
6Qwen 3 8B65.5
7DeepSeek R1 Distill Qwen 14B63.2
8DeepSeek-R1-Distill-Qwen-7B49.4
9Qwen 3 30B A3B44.5
10DeepSeek R1 Distill Llama 8B41.4
11DeepSeek R1 Distill Qwen 1.5B34

Interactive version: theaggregate.ai/benchmark?slug=quanteval-knowledge-qa-no-cot · How It Works · Data refreshed daily, snapshot 2026-09-25.