Polar (US): leaderboard

Metric: Total ICAT (0-100), the StereoSet idealized context association score adapted to political bias: the language-model score LMS (share of items where the unrelated continuation is not chosen) times the neutrality score NS (how close the preference between the left or progressive and the right or conservative continuation is to balance, averaged over issue categories); options chosen by length-normalized log-likelihood in LM Evaluation Harness; higher means more balanced and competent; mean of the economic and sociocultural axes, U.S. political statements in English. Source: arxiv.org. Saturation forecast: Around December 2026. 38 models tracked.

Top models

#ModelScore
1Llama 3.2 1B79.27
2Qwen 3 0.6B78.72
3Qwen3-0.6B-Base78.69
4Qwen3-8B-Base77.8
5Qwen3-1.7B-Base77.75
6Llama 3.2 1B Instruct76.97
7Llama 3.2 3B76.91
8Qwen 3 8B76.15
9Midm-2.0-Base-Instruct76.12
10Llama 3.2 3B Instruct76.1
11Qwen 3 1.7B76.03
12Qwen3-4B-Base75.97
13Qwen 3 4B75.36
14Qwen3-14B-Base75.14
15Mistral 7B Instruct (v0.3)74.89

Interactive version: theaggregate.ai/benchmark?slug=polar-us · How It Works · Data refreshed daily, snapshot 2026-09-29.