BALSAM - Text Manipulation: leaderboard

Metric: Overall score (0-100, LLM-judged generation and multiple choice). Source: benchmarks.ksaa.gov.sa. 29 models tracked.

Top models

#ModelScore
1Gemma 4 31B (IT)74.07
2GLM-5 (Thinking)65.08
3Command-R+64.69
4Command A 03 202564.34
5GPT-5.264.32
6Grok 4.1 Fast (Reasoning)64.12
7Kimi K2 Instruct (0905)63.73
8AceGPT-v2-8B-Chat62.26
9DeepSeek V3.261.92
10DeepSeek V4 Pro (Reasoning)61.82
11Fanar-C-2-27B60.57
12GLM-4.7 (Reasoning)60.19
13Llama 3.3 70B Instruct59.49
14Qwen 3 235B A22B59.32
15Llama 4 Maverick Instruct57.96

Interactive version: theaggregate.ai/benchmark?slug=balsam-text-manipulation · How It Works · Data refreshed daily, snapshot 2026-09-19.