Polar (South Korea): leaderboard

Metric: Total ICAT (0-100), the StereoSet idealized context association score adapted to political bias: the language-model score LMS (share of items where the unrelated continuation is not chosen) times the neutrality score NS (how close the preference between the left or progressive and the right or conservative continuation is to balance, averaged over issue categories); options chosen by length-normalized log-likelihood in LM Evaluation Harness; higher means more balanced and competent; mean of the economic and sociocultural axes, South Korean political statements in Korean. Source: arxiv.org. Saturation forecast: Estimated already saturated. 38 models tracked.

Top models

#ModelScore
1Llama 3.1 70B91.04
2Midm-2.0-Mini-Instruct90.99
3Mistral Small 390.77
4Qwen 3 14B90.75
5Qwen3-14B-Base90.45
6Mistral-Small-24B-Base-250189.4
7Llama 3.1 70B Instruct88.98
8A.X-4.0-Light88.98
9Qwen 3 8B88.7
10Midm-2.0-Base-Instruct88.17
11Qwen3-8B-Base87.9
12Qwen 3 32B87.67
13A.X 4.087.09
14Llama 3.1 8B85.83
15Qwen3-1.7B-Base84.93

Interactive version: theaggregate.ai/benchmark?slug=polar-south-korea · How It Works · Data refreshed daily, snapshot 2026-09-29.