Palmyra X5: benchmark results
Writer's Palmyra X5 enterprise model with a 1M-token context. Provider: Writer. Released 2025-04-28. Access: API.
Unified ELO 1618 ± 1, rank #216 of 1392 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Long Context - OpenAI MRCR | 25.62 | MRCR Accuracy | 100 |
| HELM Safety SimpleSafetyTests | 100 | LM Evaluated Safety score (%) | 86 |
| HELM Capabilities - GPQA | 66.14 | COT correct | 82 |
| Wolfram LLM Benchmarking Project | 56 | Correct Functionality (%) | 81.6 |
| HELM Capabilities - MMLU-Pro | 80.4 | COT correct | 76 |
| HELM Long Context - RULER HotPotQA | 57 | RULER String Match | 75 |
| HELM AIR-Bench | 78.2 | Refusal Rate (%) | 69.8 |
| HELM Capabilities - Omni-MATH | 41.45 | Acc | 68 |
| HELM Safety Anthropic Red Team | 99.3 | LM Evaluated Safety score (%) | 66.3 |
| HELM Long Context - InfiniteBench En.MC | 87 | EM | 65 |
| HELM Long Context - RULER SQuAD | 78 | RULER String Match | 65 |
| HELM Safety BBQ | 94.8 | BBQ accuracy (%) | 60.5 |
Interactive version: theaggregate.ai/model?slug=palmyra-x5 · How It Works · Data refreshed daily, snapshot 2026-09-05.