Palmyra X5 — benchmark results

Writer's Palmyra X5 enterprise model with a 1M-token context. Provider: Writer. Released 2025-04-28. Access: API.

Unified ELO 1591 ± 16, rank #501 of 1776 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Long Context - OpenAI MRCR25.62MRCR Accuracy100
HELM Safety SimpleSafetyTests100LM Evaluated Safety score (%)86
Wolfram LLM Benchmarking Project56Correct Functionality (%)83.2
HELM Capabilities - GPQA66.14COT correct82
HELM Capabilities - MMLU-Pro80.4COT correct76
HELM Long Context - RULER HotPotQA57RULER String Match75
HELM AIR-Bench78.2Refusal Rate (%)69.8
HELM Capabilities - Omni-MATH41.45Acc68
HELM Safety Anthropic Red Team99.3LM Evaluated Safety score (%)66.3
HELM Long Context - InfiniteBench En.MC87EM65
HELM Long Context - RULER SQuAD78RULER String Match65
HELM Safety BBQ94.8BBQ accuracy (%)60.5

Interactive version: theaggregate.ai/model?slug=palmyra-x5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.