Palmyra X V3 (72B) — benchmark results
Writer's 72B enterprise API model, the top non-OpenAI performer on Stanford's HELM Lite leaderboard at its debut. Provider: Writer. Released 2024-01-01. Access: API.
Unified ELO 1458 ± 45, rank #962 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM WMT 2014 | 26.19 | BLEU-4 (%) | 100 |
| HELM v2 Lite - LegalBench | 70.66 | Exact Match (%) | 95 |
| HELM v2 Lite - MedQA | 76.92 | Exact Match (%) | 87.5 |
| HELM Lite | 72.34 | Mean win rate (self-reported) | 79.2 |
| HELM (Stanford) | 67.92 | Mean Win Rate (%) | 72.2 |
| HELM NaturalQuestions (Closed) | 40.72 | F1 (%) | 65.6 |
| HELM NaturalQuestions (Open) | 68.52 | F1 (%) | 41.1 |
| HELM NarrativeQA | 70.59 | F1 (%) | 31.1 |
| HELM v2 Lite - HumanEval (Code) | 2 | Pass@1 (%) | 29.4 |
Interactive version: theaggregate.ai/model?slug=palmyra-x-v3-72b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.