Llama-3.1-Swallow-8B-Instruct-v0.5: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1489 ± 23, rank #788 of 1607 rated models, from 49 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
YOMI-Bench - Rhyme Generation84.8Accuracy (%; generated hiragana word whose vowel sequence ex100
Swallow - Post-trained Japanese - WMT20 En-Ja24.9BLEU70.1
JP-TL-Bench7.42LT Score (0-10; both directions, anchor-set Bradley-Terry fi68.4
Swallow - Post-trained Japanese - WMT20 Ja-En22.2BLEU65.7
JP-TL-Bench - English to Japanese - Easy8.91LT Score (0-10; Easy items only; anchored pairwise Bradley-T61.9
JP-TL-Bench - English to Japanese8.87LT Score (0-10; anchored pairwise Bradley-Terry vs 20 Base S60.3
JP-TL-Bench - English to Japanese - Hard9LT Score (0-10; Hard items only; anchored pairwise Bradley-T60.3
Swallow - English MT-Bench - Humanities69.1Judge Score (normalized, %)46.3
Swallow - Post-trained Japanese - JamC-QA49.6Accuracy (%)46.3
Swallow - Japanese MT-Bench - Humanities62.1Judge Score (normalized, %)44.8
YOMI-Bench - Kanji Reading Prediction (Single Reading)82.19Accuracy (%; exact match of the generated reading of a word,44.4
YOMI-Bench - Kanji Reading QA (Multiple Readings)77.82Accuracy (%; yes/no judgement of whether a given reading is 44.4

Interactive version: theaggregate.ai/model?slug=llama-3-1-swallow-8b-instruct-v0-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.