Bengali Dialect QA Consistency: leaderboard

Metric: Macro-average over the nine dialects of the Dialect consistency score (0-10): a Gemini 2.5 Flash judge (human review of low-confidence judgments) compares the model's answer to each human-corrected dialectal question with its answer to the standard Bengali original, rating dialect comprehension, factual correctness, completeness, clarity and length on weighted Likert scales; responses not in Bengali script or refusals score 0; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 3 27B (IT)8.71#509
2Llama 3.3 70B Instruct8.55#520
3GPT-OSS-120B8.2#330
4Gemma 3 12B (IT)8.13#655
5Gemma 3n E4B (IT)7.77#736
6gemma-3n-E2B-it7.52#913

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=bengali-dialect-qa-consistency · How It Works · Data refreshed daily, snapshot 2026-10-11.