Devstral Small: benchmark results
Provider: Mistral. Released 2025-05-01. Access: Open.
Unified ELO 1543 ± 45, rank #871 of 2656 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Continuation) | 487 | Elo | 31.4 |
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 84 | Questions answered with a valid query on the first attempt ( | 23.4 |
| THOR Finding Triage - Critical Miss Rate | 5.5 | True positives suppressed as false positives (%, lower is be | 20 |
| Tinybird AI SQL Benchmark - Exactness | 39.58 | Result exactness vs human reference queries (0-100) | 18.9 |
| Tinybird AI SQL Benchmark - Success Rate | 92 | Questions answered with a valid query within 3 retries (%) | 16.8 |
| THOR Finding Triage - Balanced OTS | 47.6 | Class-balanced operational triage score (%) | 16.2 |
| AI Chess Leaderboard (Reasoning) | 506 | Elo | 14.2 |
| Kagi LLM Benchmark | 34.2 | Accuracy (%) | 13.7 |
| THOR Finding Triage - CW% | 44.1 | Confidence-weighted classification score (%) | 13.7 |
| THOR Finding Triage - False Review Load | 71.1 | False positives not suppressed (%, lower is better) | 10 |
| LisanBench | 0 | Mean Path Length / Current Maximum | 6.5 |
Interactive version: theaggregate.ai/model?slug=devstral-small · How It Works · Data refreshed daily, snapshot 2026-09-19.