BALSAM - Information Extraction: leaderboard

Metric: Overall score (0-100, LLM-judged generation and multiple choice). Source: benchmarks.ksaa.gov.sa. 29 models tracked.

Top models

#ModelScore
1Gemma 4 31B (IT)57.49
2GLM-5 (Thinking)53.79
3GLM-4.7 (Reasoning)53.15
4Grok 4.1 Fast (Reasoning)52.84
5GPT-5.251.1
6Llama 4 Maverick Instruct50.66
7Kimi K2 Instruct (0905)50.36
8DeepSeek V3.250.07
9Command A 03 202549.92
10Fanar-C-2-27B49.34
11Qwen 3 235B A22B49.32
12AceGPT-v2-8B-Chat49.27
13Llama 3.3 70B Instruct49.17
14DeepSeek V4 Pro (Reasoning)47.59
15SILMA-9B-Instruct-v1.046.9

Interactive version: theaggregate.ai/benchmark?slug=balsam-information-extraction · How It Works · Data refreshed daily, snapshot 2026-09-19.