DiBiMT — leaderboard

Divide and Benchmark in Machine Translation: word sense disambiguation benchmark for MT systems across multiple language pairs.

Metric: Avg Accuracy (%). Source: nlp.uniroma1.it. Status: saturated. 26 models tracked.

Top models

#ModelScore
1GPT-471.09
2Gemma 2 9B69.8
3GPT-3.5 Turbo67.3
4Llama 3 8B55.45
5Mistral 7B48.31
6gemma-7B46.79
7Llama 2 7B46.05
8gemma-2B40.53

Interactive version: theaggregate.ai/benchmark?slug=dibimt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.