Devstral Medium — benchmark results

Mistral's API-only mid-tier agentic coding model (July 2025, built with All Hands AI) scoring 61.6% on SWE-Bench Verified at Medium-3 pricing. Provider: Mistral. Released 2025-07-10. Access: API.

Unified ELO 1488 ± 35, rank #848 of 1776 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Omniscience-31.62Score61.8
AA Omniscience - Humanities & Social Sciences22.1Accuracy (%)60
AA Omniscience - Health20.1Accuracy (%)54.9
AA Omniscience - Law11.5Accuracy (%)53.2
AA Omniscience - Software Engineering (SWE) - Java17Accuracy (%)51
AA-Omniscience Accuracy19.05Accuracy (%)49.1
AA Omniscience - Software Engineering (SWE) - Julia12Accuracy (%)47.9
AA Omniscience - Software Engineering (SWE) - HTML30Accuracy (%)47.8
AA Omniscience - Business16Accuracy (%)47.1
AA Omniscience - Software Engineering (SWE) - C34Accuracy (%)46
AA Omniscience - Software Engineering (SWE) - TypeScript20Accuracy (%)46
AA Omniscience - Software Engineering (SWE) - R10Accuracy (%)45.4

Interactive version: theaggregate.ai/model?slug=devstral-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.