Open CoT - LogiQA — leaderboard
Metric: CoT Gain (%). Source: huggingface.co. 132 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 Distill Qwen 14B | 13.74 |
| 2 | DeepSeek R1 Distill Llama 8B | 13.26 |
| 3 | DeepSeek R1 Distill Llama 70B | 12.62 |
| 4 | NeuralLLaMa-3-8B-DT-v0.1 | 11.34 |
| 5 | DeepSeek-R1-Distill-Qwen-7B | 10.38 |
| 6 | NeuralLLaMa-3-8B-ORPO-v0.3 | 10.22 |
| 7 | DeepSeek R1 Distill Qwen 32B | 10.06 |
| 8 | Qwen 2.5 3B Instruct | 9.42 |
| 9 | Daredevil-8B-abliterated | 9.42 |
| 10 | Llama 3 70B Instruct | 9.11 |
| 11 | Phi-3.5-MoE-instruct | 8.47 |
| 12 | Llama 3 8B Instruct | 8.15 |
| 13 | internlm2-chat-20B | 8.15 |
| 14 | Hermes 3 - Llama-3.1 70B | 7.99 |
| 15 | Configurable-Yi-1.5-9B-Chat | 7.99 |
Interactive version: theaggregate.ai/benchmark?slug=open-cot-logiqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.