Palmyra X (43B): benchmark results

Provider: Writer. Released 2023-06-11. Access: API.

Unified ELO 1418 ± 28, rank #1764 of 2656 rated models, from 27 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - BoolQ89.63Exact Match (%)100
HELM Classic - Entity Matching96.59Exact Match (%)100
HELM Classic - GSM8K63.3Exact Match (%)100
HELM Classic - MMLU60.91Exact Match (%)100
HELM Classic - TruthfulQA61.57Exact Match (%)100
RealToxicityPrompts0.01Toxic fraction100
HELM Classic - bAbI70.24Exact Match (%)98.6
HELM Classic - Synthetic Reasoning Abstract50.36Exact Match (%)97.1
HELM Classic - NaturalQuestions Closed Book41.27F1 (%)97
HELM Classic - BBQ78.13Exact Match (%)95.1
HELM Classic - LegalSupport62.3Exact Match (%)94.1
HELM Classic - QuAC47.29F1 (%)93.8

Interactive version: theaggregate.ai/model?slug=palmyra-x-43b · How It Works · Data refreshed daily, snapshot 2026-09-19.