Palmyra X (43B): benchmark results
Provider: Writer. Released 2023-06-11. Access: API.
Unified ELO 1418 ± 28, rank #1764 of 2656 rated models, from 27 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - BoolQ | 89.63 | Exact Match (%) | 100 |
| HELM Classic - Entity Matching | 96.59 | Exact Match (%) | 100 |
| HELM Classic - GSM8K | 63.3 | Exact Match (%) | 100 |
| HELM Classic - MMLU | 60.91 | Exact Match (%) | 100 |
| HELM Classic - TruthfulQA | 61.57 | Exact Match (%) | 100 |
| RealToxicityPrompts | 0.01 | Toxic fraction | 100 |
| HELM Classic - bAbI | 70.24 | Exact Match (%) | 98.6 |
| HELM Classic - Synthetic Reasoning Abstract | 50.36 | Exact Match (%) | 97.1 |
| HELM Classic - NaturalQuestions Closed Book | 41.27 | F1 (%) | 97 |
| HELM Classic - BBQ | 78.13 | Exact Match (%) | 95.1 |
| HELM Classic - LegalSupport | 62.3 | Exact Match (%) | 94.1 |
| HELM Classic - QuAC | 47.29 | F1 (%) | 93.8 |
Interactive version: theaggregate.ai/model?slug=palmyra-x-43b · How It Works · Data refreshed daily, snapshot 2026-09-19.