KOR-Bench - Operation: leaderboard

Metric: Accuracy (%; zero-shot, 250 questions). Source: arxiv.org. Saturation forecast: Estimated already saturated. 38 models tracked.

Top models

#ModelScore
1Qwen 2.5 32B Instruct93.2
2GPT-4 Turbo90.4
3O1 Preview (2024-09-12)88.8
4Claude 3.5 Sonnet (20240620)88.4
5Llama 3.1 405B Instruct87.82
6Mistral Large 2 (Jul)86.8
7GPT-4o (2024-05-13)86
8Llama 3.1 70B Instruct84.8
9Qwen 2.5 14B Instruct84.4
10Yi Large84
11Qwen 2.5 72B Instruct83.6
12O1 Mini (2024-09-12)82.8
13Llama 3 70B Instruct82.4
14Gemini 1.5 Pro81.6
15Yi 1.5 34B Chat79.6

Interactive version: theaggregate.ai/benchmark?slug=kor-bench-operation · How It Works · Data refreshed daily, snapshot 2026-09-24.