Translation en->Set1 spBleu: leaderboard

Translation evaluation using spBLEU (SentencePiece BLEU), a BLEU metric computed over text tokenized with a language-agnostic SentencePiece subword model. Introduced in the FLORES-101 evaluation benchmark for low-resource and multilingual machine translation.

Metric: en→Set1 spBLEU (self-reported). Source: benchmarklist.com. 13 models tracked.

Top models

#ModelScore
1Nova Pro43.4
2GPT-4o43.1
3Gemini 1.5 Pro (002)43
4Claude 3.5 Sonnet42.5
5Nova Lite41.5
6GPT-4o Mini41.1
7Nova Micro40.2
8Gemini 1.5 Flash (002)40
9Claude 3.5 Haiku40
10Llama 3.2 90B39.7
11Gemini 1.5 Flash-8B (001)38.2
12Llama 3.1 8B32.7

Interactive version: theaggregate.ai/benchmark?slug=translation-en-over-set1-spbleu · How It Works · Data refreshed daily, snapshot 2026-09-05.