KOR-Bench - Logic: leaderboard

Metric: Accuracy (%; zero-shot, 250 questions). Source: arxiv.org. Saturation forecast: Estimated already saturated. 38 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20240620)67.2
2O1 Preview (2024-09-12)63.2
3O1 Mini (2024-09-12)61.2
4Llama 3.1 405B Instruct56.8
5Qwen 2.5 32B Instruct56.8
6GPT-4 Turbo54
7Qwen 2.5 72B Instruct53.2
8GPT-4o (2024-05-13)52.4
9Mistral Large 2 (Jul)51.2
10Qwen 2.5 14B Instruct50
11Llama 3.1 70B Instruct49.2
12Gemma 2 27B (IT)49.2
13DeepSeek V2.548
14Yi Large47.6
15Llama 3 70B Instruct46.4

Interactive version: theaggregate.ai/benchmark?slug=kor-bench-logic · How It Works · Data refreshed daily, snapshot 2026-09-24.