Devstral Small: benchmark results

Provider: Mistral. Released 2025-05-01. Access: Open.

Unified ELO 1543 ± 45, rank #871 of 2656 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Chess Leaderboard (Continuation)487Elo31.4
Tinybird AI SQL Benchmark - First-Attempt Success Rate84Questions answered with a valid query on the first attempt (23.4
THOR Finding Triage - Critical Miss Rate5.5True positives suppressed as false positives (%, lower is be20
Tinybird AI SQL Benchmark - Exactness39.58Result exactness vs human reference queries (0-100)18.9
Tinybird AI SQL Benchmark - Success Rate92Questions answered with a valid query within 3 retries (%)16.8
THOR Finding Triage - Balanced OTS47.6Class-balanced operational triage score (%)16.2
AI Chess Leaderboard (Reasoning)506Elo14.2
Kagi LLM Benchmark34.2Accuracy (%)13.7
THOR Finding Triage - CW%44.1Confidence-weighted classification score (%)13.7
THOR Finding Triage - False Review Load71.1False positives not suppressed (%, lower is better)10
LisanBench0Mean Path Length / Current Maximum6.5

Interactive version: theaggregate.ai/model?slug=devstral-small · How It Works · Data refreshed daily, snapshot 2026-09-19.