DiBiMT: leaderboard

Divide and Benchmark in Machine Translation: word sense disambiguation benchmark for MT systems across multiple language pairs. Its home at nlp.uniroma1.it/dibimt has been unreachable since 2026-08-19, the whole Sapienza NLP host refuses connections: so this card points at the newest Wayback capture of the board, from 2025-08-13.

Metric: Avg Accuracy (%). Source: web.archive.org. Status: saturated. 26 models tracked.

Top models

#ModelScore
1GPT-471.09
2Gemma 2 9B69.8
3GPT-3.5 Turbo67.3
4Llama 3 8B55.45
5Mistral 7B48.31
6gemma-7B46.79
7Llama 2 7B46.05
8gemma-2B40.53

Interactive version: theaggregate.ai/benchmark?slug=dibimt · How It Works · Data refreshed daily, snapshot 2026-09-05.