Palmyra Med — benchmark results
Writer's 70B medical-domain specialist tuned on biomedical data, released July 2024 alongside Palmyra Fin. Provider: Writer. Released 2024-07-30. Access: API.
Unified ELO 1429 ± 23, rank #1089 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety XSTest | 96.3 | LM Evaluated Safety score (%) | 55.8 |
| HELM Safety SimpleSafetyTests | 98.5 | LM Evaluated Safety score (%) | 46.5 |
| HELM AIR-Bench | 57.8 | Refusal Rate (%) | 31.4 |
| HELM Safety Anthropic Red Team | 97.8 | LM Evaluated Safety score (%) | 27.9 |
| HELM Capabilities - GPQA | 36.77 | COT correct | 23 |
| HELM Capabilities - IFEval | 76.74 | IFEval Strict Acc | 22 |
| HELM Safety | 85.7 | Mean score (self-reported) | 21.5 |
| HELM Safety BBQ | 80.1 | BBQ accuracy (%) | 19.8 |
| HELM Safety HarmBench | 55.6 | LM Evaluated Safety score (%) | 18.6 |
| HELM Capabilities - MMLU-Pro | 41.1 | COT correct | 14 |
| HELM Capabilities - WildBench | 67.58 | WB Score | 12 |
| HELM Capabilities - Omni-MATH | 15.57 | Acc | 11 |
Interactive version: theaggregate.ai/model?slug=palmyra-med · How the rankings work · Data refreshed daily, snapshot 2026-07-22.