Devstral Small (May '25): benchmark results

Provider: Mistral. Released 2025-05-01. Access: Open.

Unified ELO 1490 ± 1, rank #911 of 1761 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA TAU-2 Bench38.01Accuracy (%)46
AA Omniscience - Humanities & Social Sciences17.8Accuracy (%)41.3
AA Omniscience - Law10Accuracy (%)40.6
Artificial Analysis Intelligence Index6.04Intelligence Index37.3
AA Omniscience - Health17.3Accuracy (%)37.2
AA Terminal-Bench Hard6.06Accuracy (%)33.9
AA Long Context Reasoning32Accuracy (%)33.6
AA Omniscience - Software Engineering (SWE)20.2Accuracy (%)33.1
AA-Omniscience Accuracy15.87Accuracy (%)29.4
AA CritPt0Accuracy (%)24.4
AA Omniscience-56.78Score23.2
AA Omniscience - Science, Engineering & Mathematics18.8Accuracy (%)20.6

Interactive version: theaggregate.ai/model?slug=devstral-small-may-25 · How It Works · Data refreshed daily, snapshot 2026-09-05.