K-MetBench — leaderboard

Expert meteorology benchmark over 1,774 Korean National Meteorological Engineer Examination questions, including reasoning, geo-cultural, text-only, and multimodal subsets.

Metric: Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 59 models tracked.

Top models

#ModelScore
1GPT-5.287.8
2Qwen 3 VL 235B A22B (Thinking)84.4
3Qwen 3.5 27B83
4Qwen 3.6 35B A3B82.9
5Qwen 3 VL 32B (Thinking)78.6
6command-a-reasoning-08-202577.8
7GPT-OSS-120B77.3
8Qwen 3 30B A3B 2507 (Thinking)76.7
9EXAONE 4.5 33B75.9
10Qwen 3.5 9B74.9
11Qwen 3 VL 30B A3B (Thinking)74.9
12Qwen 3 14B73.7
13Qwen 3 VL 235B A22B Instruct72.4
14Qwen 3 VL 8B (Thinking)71.7
15GPT-OSS-20B71.5

Interactive version: theaggregate.ai/benchmark?slug=k-metbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.