Llama-PLLuM-70B-chat: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1556 ± 23, rank #813 of 2656 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Polish EQ-Bench | 72.99 | EQ-Bench Score | 93.1 |
| LLMZSZL Leaderboard | 64.42 | Score | 91.8 |
| MT-Bench PL - Extraction | 9.45 | Judge Score (0-10) | 77.6 |
| MT-Bench PL - Writing | 8.05 | Judge Score (0-10) | 57.1 |
| CPTU Bench | 3.53 | Average Score (1-5) | 53.5 |
| MT-Bench PL - Overall | 6.75 | Judge Score (0-10) | 51 |
| MT-Bench PL - Reasoning | 5.2 | Judge Score (0-10) | 45.9 |
| MT-Bench PL - Coding | 4.8 | Judge Score (0-10) | 44.9 |
| PLCC - Culture & Tradition | 64 | Accuracy (%) | 44.7 |
| PLCC - History | 74 | Accuracy (%) | 44.2 |
| MT-Bench PL - STEM | 8.2 | Judge Score (0-10) | 39.8 |
| MT-Bench PL - Humanities | 8.8 | Judge Score (0-10) | 38.8 |
Interactive version: theaggregate.ai/model?slug=llama-pllum-70b-chat · How It Works · Data refreshed daily, snapshot 2026-09-19.