Translation Set1->en spBleu: leaderboard

spBLEU (SentencePiece BLEU) evaluation metric for machine translation quality assessment, using language-agnostic SentencePiece tokenization with BLEU scoring. Part of the FLORES-101 evaluation benchmark for low-resource and multilingual machine translation.

Metric: Set1→en spBLEU (self-reported). Source: benchmarklist.com. 13 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro (002)45.6
2Nova Pro44.4
3GPT-4o43.9
4Llama 3.2 90B43.7
5Claude 3.5 Sonnet43.5
6Nova Lite43.1
7Gemini 1.5 Flash (002)42.9
8Nova Micro42.6
9GPT-4o Mini41.9
10Gemini 1.5 Flash-8B (001)41.4
11Claude 3.5 Haiku40.2
12Llama 3.1 8B36.5

Interactive version: theaggregate.ai/benchmark?slug=translation-set1-over-en-spbleu · How It Works · Data refreshed daily, snapshot 2026-09-05.