SALMONN-13B: benchmark results
Provider: Other.
Unified ELO 1409 ± 24, rank #1124 of 1605 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Afrispeech Semantics - Consistency (AfriSpeech-General) | 90.52 | Macro-F1 (%; binary consistent versus inconsistent judgments | 80 |
| Afrispeech Semantics - Consistency (AfriSpeech-Medical) | 77.49 | Macro-F1 (%; binary consistent versus inconsistent judgments | 80 |
| CASU - Entity Extraction | 50.96 | Accuracy (%; four-option multiple-choice questions on semi-s | 66.7 |
| Afrispeech Semantics - Consistency (AfriSpeech-200) | 73.75 | Macro-F1 (%; binary consistent versus inconsistent judgments | 60 |
| CASU - Scene Description Event Recall | 36 | Event match (%; an LLM judge counts the share of the scene s | 50 |
| CASU - Counterfactual Reasoning | 46.95 | Accuracy (%; four-option multiple-choice questions on semi-s | 41.7 |
| Afrispeech Semantics - Entailment (AfriSpeech-200) | 28.14 | Macro-F1 (%; three-way entailment, neutral or contradiction | 33.3 |
| Afrispeech Semantics - Plausibility (AfriSpeech-200) | 59.99 | Macro-F1 (%; plausible-but-unsupported versus implausible ju | 33.3 |
| Afrispeech Semantics - Plausibility (AfriSpeech-General) | 63.3 | Macro-F1 (%; plausible-but-unsupported versus implausible ju | 33.3 |
| CASU - Contextual Reasoning | 43.21 | Accuracy (%; four-option multiple-choice questions on semi-s | 33.3 |
| CASU - Role Inference | 39.73 | Accuracy (%; four-option multiple-choice questions on semi-s | 33.3 |
| FinBen - FinNum | 0 | Normalized Score | 30 |
Interactive version: theaggregate.ai/model?slug=salmonn-13b · How It Works · Data refreshed daily, snapshot 2026-09-26.