Devstral Medium: benchmark results
Mistral's API-only mid-tier agentic coding model (July 2025, built with All Hands AI) scoring 61.6% on SWE-Bench Verified at Medium-3 pricing. Provider: Mistral. Released 2025-07-10. Access: API.
Unified ELO 1526 ± 1, rank #566 of 1392 rated models, from 49 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - Dart | 20 | Accuracy (%) | 59.4 |
| AA Omniscience - Software Engineering (SWE) - Julia | 12 | Accuracy (%) | 58 |
| AA Omniscience - Software Engineering (SWE) - Java | 16 | Accuracy (%) | 57 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 18 | Accuracy (%) | 55.3 |
| AA Omniscience | -31.57 | Score | 55.2 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 18.89 | Accuracy (%) | 53.8 |
| AA Omniscience - Humanities & Social Sciences | 21.2 | Accuracy (%) | 53.7 |
| AA Omniscience - Software Engineering (SWE) - R | 10 | Accuracy (%) | 53.4 |
| AA Omniscience - Software Engineering (SWE) - C | 32 | Accuracy (%) | 52.9 |
| AA Omniscience - Software Engineering (SWE) - HTML | 30 | Accuracy (%) | 51.7 |
| AA Omniscience - Health | 20.6 | Accuracy (%) | 51.3 |
| AA Omniscience - Software Engineering (SWE) - Python | 18 | Accuracy (%) | 51.2 |
Interactive version: theaggregate.ai/model?slug=devstral-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.