Palmyra-X-004 — benchmark results

Writer's enterprise frontier model emphasizing tool calling and built-in RAG for AI agents and workflows (October 2024). Provider: Writer. Released 2024-10-09. Access: API.

Unified ELO 1548 ± 22, rank #624 of 1776 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Safety XSTest98.4LM Evaluated Safety score (%)92.4
HELM Lite85.05Mean win rate (self-reported)92.2
HELM NaturalQuestions (Closed)45.66F1 (%)87.8
HELM Safety SimpleSafetyTests100LM Evaluated Safety score (%)86
HELM (Stanford)80.82Mean Win Rate (%)84.4
HELM NarrativeQA77.26F1 (%)84.4
HELM Capabilities - IFEval87.25IFEval Strict Acc82
HELM NaturalQuestions (Open)75.38F1 (%)78.9
HELM Safety BBQ95.5BBQ accuracy (%)70.9
HELM Safety93.3Mean score (self-reported)64
HELM WMT 201420.33BLEU-4 (%)62.8
HELM Capabilities - WildBench80.2WB Score62

Interactive version: theaggregate.ai/model?slug=palmyra-x-004 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.