Open CoT Leaderboard — leaderboard
Measures chain-of-thought reasoning effectiveness: accuracy gain from using CoT prompting across LogiQA and LSAT logical reasoning tasks.
Metric: Average CoT Gain (%). Source: huggingface.co. Status: saturated. 133 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 Distill Qwen 14B | 17.65 |
| 2 | DeepSeek R1 Distill Qwen 32B | 16.92 |
| 3 | Llama 3.1 8B Instruct | 16.29 |
| 4 | internlm2-chat-20B | 16.11 |
| 5 | DeepSeek R1 Distill Llama 8B | 16.1 |
| 6 | DeepSeek R1 Distill Llama 70B | 15.3 |
| 7 | Llama 3 8B Instruct | 14.75 |
| 8 | NeuralLLaMa-3-8B-ORPO-v0.3 | 14.52 |
| 9 | Daredevil-8B-abliterated | 14.25 |
| 10 | Llama 3 70B Instruct | 13.98 |
| 11 | Llama-3.1-SuperNova-Lite | 13.97 |
| 12 | Llama 3.1 Nemotron 70B Instruct | 13.87 |
| 13 | Yi 1.5 34B Chat | 13.81 |
| 14 | Llama 3.1 70B Instruct | 13.39 |
| 15 | NeuralLLaMa-3-8B-DT-v0.1 | 13.33 |
Interactive version: theaggregate.ai/benchmark?slug=open-cot-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.