Open CoT Leaderboard — leaderboard

Measures chain-of-thought reasoning effectiveness: accuracy gain from using CoT prompting across LogiQA and LSAT logical reasoning tasks.

Metric: Average CoT Gain (%). Source: huggingface.co. Status: saturated. 133 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Qwen 14B17.65
2DeepSeek R1 Distill Qwen 32B16.92
3Llama 3.1 8B Instruct16.29
4internlm2-chat-20B16.11
5DeepSeek R1 Distill Llama 8B16.1
6DeepSeek R1 Distill Llama 70B15.3
7Llama 3 8B Instruct14.75
8NeuralLLaMa-3-8B-ORPO-v0.314.52
9Daredevil-8B-abliterated14.25
10Llama 3 70B Instruct13.98
11Llama-3.1-SuperNova-Lite13.97
12Llama 3.1 Nemotron 70B Instruct13.87
13Yi 1.5 34B Chat13.81
14Llama 3.1 70B Instruct13.39
15NeuralLLaMa-3-8B-DT-v0.113.33

Interactive version: theaggregate.ai/benchmark?slug=open-cot-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.