DialSeg-Ar - Code-Switched Podcasts: leaderboard

Metric: WindowDiff (0-1, lower is better; share of sliding windows over the utterance sequence whose predicted boundary count differs from the gold count; test split of Gulf Arabic and English code-switched podcast transcripts; LLMs prompted zero-shot, DialSeg-Ar-Gemma3-4B fine-tuned on the training split). Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1Gemma 3 4B (IT)0.67
2Fanar-1-9B-Instruct0.7
3ALLaM-7B-Instruct-preview0.81

Interactive version: theaggregate.ai/benchmark?slug=dialseg-ar-code-switched-podcasts · How It Works · Data refreshed daily, snapshot 2026-09-26.