DialSeg-Ar - OPUS News: leaderboard

Metric: WindowDiff (0-1, lower is better; share of sliding windows over the utterance sequence whose predicted boundary count differs from the gold count; test split of OPUS News Commentary in Modern Standard Arabic; LLMs prompted zero-shot, DialSeg-Ar-Gemma3-4B fine-tuned on the training split). Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1Fanar-1-9B-Instruct0.57
2ALLaM-7B-Instruct-preview0.65
3Gemma 3 4B (IT)0.75

Interactive version: theaggregate.ai/benchmark?slug=dialseg-ar-opus-news · How It Works · Data refreshed daily, snapshot 2026-09-26.