SALMONN-13B: benchmark results

Provider: Other.

Unified ELO 1409 ± 24, rank #1124 of 1605 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Afrispeech Semantics - Consistency (AfriSpeech-General)90.52Macro-F1 (%; binary consistent versus inconsistent judgments80
Afrispeech Semantics - Consistency (AfriSpeech-Medical)77.49Macro-F1 (%; binary consistent versus inconsistent judgments80
CASU - Entity Extraction50.96Accuracy (%; four-option multiple-choice questions on semi-s66.7
Afrispeech Semantics - Consistency (AfriSpeech-200)73.75Macro-F1 (%; binary consistent versus inconsistent judgments60
CASU - Scene Description Event Recall36Event match (%; an LLM judge counts the share of the scene s50
CASU - Counterfactual Reasoning46.95Accuracy (%; four-option multiple-choice questions on semi-s41.7
Afrispeech Semantics - Entailment (AfriSpeech-200)28.14Macro-F1 (%; three-way entailment, neutral or contradiction 33.3
Afrispeech Semantics - Plausibility (AfriSpeech-200)59.99Macro-F1 (%; plausible-but-unsupported versus implausible ju33.3
Afrispeech Semantics - Plausibility (AfriSpeech-General)63.3Macro-F1 (%; plausible-but-unsupported versus implausible ju33.3
CASU - Contextual Reasoning43.21Accuracy (%; four-option multiple-choice questions on semi-s33.3
CASU - Role Inference39.73Accuracy (%; four-option multiple-choice questions on semi-s33.3
FinBen - FinNum0Normalized Score30

Interactive version: theaggregate.ai/model?slug=salmonn-13b · How It Works · Data refreshed daily, snapshot 2026-09-26.