Open CoT - LogiQA 2 — leaderboard
Metric: CoT Gain (%). Source: huggingface.co. 132 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek R1 Distill Qwen 32B | 19.02 |
| 2 | DeepSeek R1 Distill Qwen 14B | 18.7 |
| 3 | DeepSeek R1 Distill Llama 8B | 18.26 |
| 4 | internlm2-chat-20B | 17.62 |
| 5 | NeuralLLaMa-3-8B-ORPO-v0.3 | 16.48 |
| 6 | Llama 3.1 8B Instruct | 15.52 |
| 7 | Phi-3.5-MoE-instruct | 15.14 |
| 8 | Llama 3 70B Instruct | 15.08 |
| 9 | SOLAR-10.7B-Instruct-v1.0 | 15.01 |
| 10 | Daredevil-8B-abliterated | 15.01 |
| 11 | NeuralLLaMa-3-8B-DT-v0.1 | 14.89 |
| 12 | Command-R+ (08-2024) | 14.76 |
| 13 | openbuddy-yi1.5-9B-v21.1-32k | 14.69 |
| 14 | DeepSeek R1 Distill Llama 70B | 14.25 |
| 15 | Llama 3.1 70B Instruct | 14.19 |
Interactive version: theaggregate.ai/benchmark?slug=open-cot-logiqa-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.