KOR-Bench - Base Models (3-shot): leaderboard

Metric: Accuracy (%; three-shot, 1,250 questions). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.

Top models

#ModelScore
1Llama 3.1 405B39.68
2Qwen 2.5 72B37.28
3Qwen 2.5 32B37.28
4Llama 3 70B35.2
5Qwen 2 72B34.32
6Llama 3.1 70B33.84
7Gemma 2 27B33.36
8Qwen 2.5 14B33.28
9Yi 1.5 34B30.08
10Yi-1.5-9B29.2
11Qwen 2.5 7B28.8
12Qwen 2 7B27.44
13Llama 3.1 8B26
14Gemma 2 9B25.52
15Llama 3 8B24.96

Interactive version: theaggregate.ai/benchmark?slug=kor-bench-base-models-3-shot · How It Works · Data refreshed daily, snapshot 2026-09-24.