BenSyc: leaderboard

Metric: Five-class conversational alignment classification macro-F1 (%) (invalidation, neutral, support, validation, escalation) on human-annotated Bengali and Banglish Reddit replies from Bangladesh and West Bengal communities, deterministic decoding (temperature 0), open models via Ollama; the paper ranks the leaderboard by this score; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Gemma 4 31B61.7
2GPT-5.4 Mini57.2
3Qwen 2.5 32B55.4
4Llama 3.3 70B54.3
5Qwen 2.5 14B44.2
6Llama 3.1 8B41.2
7Gemma 2 9B40.3
8Qwen 2.5 7B38.2
9Gemma 2 27B38.2
10Mixtral 8x7B33.2
11Mistral 7B Instruct30.5
12Llama 3.2 3B21.3

Interactive version: theaggregate.ai/benchmark?slug=bensyc · How It Works · Data refreshed daily, snapshot 2026-09-29.