Llama-3.1-Swallow-8B-Instruct-v0.5: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1489 ± 23, rank #788 of 1607 rated models, from 49 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| YOMI-Bench - Rhyme Generation | 84.8 | Accuracy (%; generated hiragana word whose vowel sequence ex | 100 |
| Swallow - Post-trained Japanese - WMT20 En-Ja | 24.9 | BLEU | 70.1 |
| JP-TL-Bench | 7.42 | LT Score (0-10; both directions, anchor-set Bradley-Terry fi | 68.4 |
| Swallow - Post-trained Japanese - WMT20 Ja-En | 22.2 | BLEU | 65.7 |
| JP-TL-Bench - English to Japanese - Easy | 8.91 | LT Score (0-10; Easy items only; anchored pairwise Bradley-T | 61.9 |
| JP-TL-Bench - English to Japanese | 8.87 | LT Score (0-10; anchored pairwise Bradley-Terry vs 20 Base S | 60.3 |
| JP-TL-Bench - English to Japanese - Hard | 9 | LT Score (0-10; Hard items only; anchored pairwise Bradley-T | 60.3 |
| Swallow - English MT-Bench - Humanities | 69.1 | Judge Score (normalized, %) | 46.3 |
| Swallow - Post-trained Japanese - JamC-QA | 49.6 | Accuracy (%) | 46.3 |
| Swallow - Japanese MT-Bench - Humanities | 62.1 | Judge Score (normalized, %) | 44.8 |
| YOMI-Bench - Kanji Reading Prediction (Single Reading) | 82.19 | Accuracy (%; exact match of the generated reading of a word, | 44.4 |
| YOMI-Bench - Kanji Reading QA (Multiple Readings) | 77.82 | Accuracy (%; yes/no judgement of whether a given reading is | 44.4 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-swallow-8b-instruct-v0-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.