BanglaSocialBench - Social Customs: leaderboard

Metric: Accuracy (%) on the four-option Bengali social-customs items (pick the pragmatically appropriate response to an everyday scenario; 392 written, 390 scored), zero-shot Bangla prompts at temperature 0, culturally grounded scenarios written and verified by native Bangla speakers; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 486.67#194
2Gemini 2.5 Flash86.41#237
3GPT-4o86.15#333
4Gemini 2.0 Flash85.13#331
5DeepSeek V3.183.33#260
6Llama 3.3 70B Instruct81.54#520
7Gemma 3 27B81.02#596
8GPT-4o Mini77.43#588
9Qwen 2.5 72B Instruct77.18#436
10Gemma 3 12B76.15#666
11Claude 3.5 Haiku71.02#553
12Llama 3 8B Instruct42.82#1115

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=banglasocialbench-social-customs · How It Works · Data refreshed daily, snapshot 2026-10-11.