Palmyra X V2 (33B): benchmark results

Writer's API-only 33B Palmyra-X V2 model (December 2023), pretrained on over 2T tokens and a strong HELM Lite performer for its size. Provider: Writer. Released 2023-12-01. Access: API.

Unified ELO 1555 ± 1, rank #421 of 1392 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM WMT 201423.88BLEU-4 (%)95.6
HELM v2 Lite - LegalBench61.37Exact Match (%)80
HELM NaturalQuestions (Open)75.15F1 (%)77.8
HELM NaturalQuestions (Closed)42.81F1 (%)74.4
HELM Lite63.98Mean win rate (self-reported)71.1
HELM (Stanford)58.89Mean Win Rate (%)64.4
HELM NarrativeQA75.25F1 (%)64.4
HELM v2 Lite - MedQA67.31Exact Match (%)57.5
HELM v2 Lite - HumanEval (Code)0Pass@1 (%)11.8

Interactive version: theaggregate.ai/model?slug=palmyra-x-v2-33b · How It Works · Data refreshed daily, snapshot 2026-09-05.