Palmyra X V2 (33B) — benchmark results
Writer's API-only 33B Palmyra-X V2 model (December 2023), pretrained on over 2T tokens and a strong HELM Lite performer for its size. Provider: Writer. Released 2023-12-01. Access: API.
Unified ELO 1410 ± 70, rank #1176 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM WMT 2014 | 23.88 | BLEU-4 (%) | 95.6 |
| HELM v2 Lite - LegalBench | 61.37 | Exact Match (%) | 80 |
| HELM NaturalQuestions (Open) | 75.15 | F1 (%) | 77.8 |
| HELM NaturalQuestions (Closed) | 42.81 | F1 (%) | 74.4 |
| HELM Lite | 63.98 | Mean win rate (self-reported) | 71.4 |
| HELM (Stanford) | 58.89 | Mean Win Rate (%) | 64.4 |
| HELM NarrativeQA | 75.25 | F1 (%) | 64.4 |
| HELM v2 Lite - MedQA | 67.31 | Exact Match (%) | 57.5 |
| HELM v2 Lite - HumanEval (Code) | 0 | Pass@1 (%) | 11.8 |
Interactive version: theaggregate.ai/model?slug=palmyra-x-v2-33b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.