TelcoAgent-Bench - Sequence Alignment (Arabic): leaderboard

Metric: Sequence alignment score x100 (0-100) on TelcoAgent-Bench (15 telecom troubleshooting intents, 49 blueprints with 30 sampled dialogues each; the agent starts from the engineer's first message and calls core and distractor network tools), Arabic dialogues: per sample, the longest common subsequence of the agent's tool calls with the gold tool path over the gold length, times a penalty for calls beyond the gold length; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Granite 3.3 8B Instruct62.3
2Qwen 3 8B60.3
3Llama 3 8B Instruct47.2
4Gemma 3 4B (IT)46.4
5Mistral 7B Instruct43.6
6Qwen 2.5 7B Instruct42.4
7granite-3.1-3B-a800m-instruct33.4
8Qwen-7B Chat30

Interactive version: theaggregate.ai/benchmark?slug=telcoagent-bench-sequence-alignment-arabic · How It Works · Data refreshed daily, snapshot 2026-10-07.