BlueBench - Translation — leaderboard

Metric: Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct41.71
2GPT-4o41.42
3GPT-4.141.22
4GPT-4.1 Nano40.29
5GPT-4.1 Mini40.16
6Mistral Large39.83
7O3 Mini37.79
8O137.44
9O4 Mini33.43
10Pixtral-12B32.54
11Llama 3.2 3B Instruct31.01
12Llama 3.2 1B Instruct23.29
13Mistral Medium 315.52

Interactive version: theaggregate.ai/benchmark?slug=bluebench-translation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.