Palmyra Fin — benchmark results
Writer's finance-specialist Palmyra model (70B, 32K context), notable as the first model to pass a CFA Level III sample exam (July 2024). Provider: Writer. Released 2024-07-30. Access: API.
Unified ELO 1561 ± 20, rank #588 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety SimpleSafetyTests | 100 | LM Evaluated Safety score (%) | 86 |
| HELM Safety Anthropic Red Team | 99.5 | LM Evaluated Safety score (%) | 76.7 |
| HELM Safety | 94.8 | Mean score (self-reported) | 71.1 |
| HELM Safety HarmBench | 84.1 | LM Evaluated Safety score (%) | 62.2 |
| HELM Safety BBQ | 94.2 | BBQ accuracy (%) | 55.8 |
| HELM Safety XSTest | 96.2 | LM Evaluated Safety score (%) | 54.7 |
| HELM AIR-Bench | 66.3 | Refusal Rate (%) | 51.2 |
| HELM Capabilities - Omni-MATH | 29.47 | Acc | 44 |
| HELM Capabilities - WildBench | 78.34 | WB Score | 40 |
| HELM Capabilities - GPQA | 42.15 | COT correct | 36 |
| HELM Capabilities - IFEval | 79.3 | IFEval Strict Acc | 34 |
| HELM Capabilities - MMLU-Pro | 59.1 | COT correct | 26 |
Interactive version: theaggregate.ai/model?slug=palmyra-fin · How the rankings work · Data refreshed daily, snapshot 2026-07-22.