Devstral Medium — benchmark results
Mistral's API-only mid-tier agentic coding model (July 2025, built with All Hands AI) scoring 61.6% on SWE-Bench Verified at Medium-3 pricing. Provider: Mistral. Released 2025-07-10. Access: API.
Unified ELO 1488 ± 35, rank #848 of 1776 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience | -31.62 | Score | 61.8 |
| AA Omniscience - Humanities & Social Sciences | 22.1 | Accuracy (%) | 60 |
| AA Omniscience - Health | 20.1 | Accuracy (%) | 54.9 |
| AA Omniscience - Law | 11.5 | Accuracy (%) | 53.2 |
| AA Omniscience - Software Engineering (SWE) - Java | 17 | Accuracy (%) | 51 |
| AA-Omniscience Accuracy | 19.05 | Accuracy (%) | 49.1 |
| AA Omniscience - Software Engineering (SWE) - Julia | 12 | Accuracy (%) | 47.9 |
| AA Omniscience - Software Engineering (SWE) - HTML | 30 | Accuracy (%) | 47.8 |
| AA Omniscience - Business | 16 | Accuracy (%) | 47.1 |
| AA Omniscience - Software Engineering (SWE) - C | 34 | Accuracy (%) | 46 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 20 | Accuracy (%) | 46 |
| AA Omniscience - Software Engineering (SWE) - R | 10 | Accuracy (%) | 45.4 |
Interactive version: theaggregate.ai/model?slug=devstral-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.