Palmyra Fin: benchmark results
Writer's finance-specialist Palmyra model (70B, 32K context), notable as the first model to pass a CFA Level III sample exam (July 2024). Provider: Writer. Released 2024-07-30. Access: API.
Unified ELO 1572 ± 1, rank #355 of 1392 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety SimpleSafetyTests | 100 | LM Evaluated Safety score (%) | 86 |
| HELM Safety Anthropic Red Team | 99.5 | LM Evaluated Safety score (%) | 76.7 |
| HELM Safety | 94.8 | Mean score (self-reported) | 71 |
| HELM Safety HarmBench | 84.1 | LM Evaluated Safety score (%) | 62.2 |
| HELM Safety BBQ | 94.2 | BBQ accuracy (%) | 55.8 |
| HELM Safety XSTest | 96.2 | LM Evaluated Safety score (%) | 54.7 |
| HELM AIR-Bench | 66.3 | Refusal Rate (%) | 51.2 |
| HELM Capabilities - Omni-MATH | 29.47 | Acc | 44 |
| HELM Capabilities - WildBench | 78.34 | WB Score | 40 |
| HELM Capabilities - GPQA | 42.15 | COT correct | 36 |
| HELM Capabilities - IFEval | 79.3 | IFEval Strict Acc | 34 |
| HELM Capabilities - MMLU-Pro | 59.1 | COT correct | 26 |
Interactive version: theaggregate.ai/model?slug=palmyra-fin · How It Works · Data refreshed daily, snapshot 2026-09-05.