Palmyra X V2 (33B) — benchmark results

Writer's API-only 33B Palmyra-X V2 model (December 2023), pretrained on over 2T tokens and a strong HELM Lite performer for its size. Provider: Writer. Released 2023-12-01. Access: API.

Unified ELO 1410 ± 70, rank #1176 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM WMT 201423.88BLEU-4 (%)95.6
HELM v2 Lite - LegalBench61.37Exact Match (%)80
HELM NaturalQuestions (Open)75.15F1 (%)77.8
HELM NaturalQuestions (Closed)42.81F1 (%)74.4
HELM Lite63.98Mean win rate (self-reported)71.4
HELM (Stanford)58.89Mean Win Rate (%)64.4
HELM NarrativeQA75.25F1 (%)64.4
HELM v2 Lite - MedQA67.31Exact Match (%)57.5
HELM v2 Lite - HumanEval (Code)0Pass@1 (%)11.8

Interactive version: theaggregate.ai/model?slug=palmyra-x-v2-33b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.