Devstral Small 2 — benchmark results

Mistral's Apache-2.0 24B agentic coding model with a 256K context, strong on SWE-bench Verified for its size (December 2025). Provider: Mistral. Released 2025-12-09. Access: Open.

Unified ELO 1533 ± 14, rank #672 of 1776 rated models, from 100 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Konkur 1404 - Art55.17Accuracy (%, text-only)94.7
Konkur 1404 - Humanities45.7Accuracy (%, text-only)89.5
YapBench449YapIndex (lower is better)85.4
Konkur 1404 - Overall43.8Accuracy (%, text-only)84.2
UGI Leaderboard45.28UGI Score81.3
Konkur 1404 - Mathematics44.35Accuracy (%, text-only)76.3
Konkur 1404 - Experimental Sciences39.13Accuracy (%, text-only)68.4
UGI - Willingness (W/10)6.8W/10 Score66.6
Konkur 1404 - Foreign Language37.5Accuracy (%, text-only)65.8
AI Chess Leaderboard (Continuation)602Elo54.5
AA Terminal-Bench Hard16.67Accuracy (%)54.2
Artificial Analysis Intelligence Index17.44Intelligence Index53.9

Interactive version: theaggregate.ai/model?slug=devstral-small-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.