maestrale-chat-v0.4-beta — benchmark results
mii-llm's Italian chat model built on Mistral-7B with continued Italian pretraining, SFT on 1.7M conversations, and DPO alignment. Provider: Other. Released 2024-06-06. Access: Open.
Unified ELO 1446 ± 8, rank #1017 of 1776 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Pinocchio Italian - Cultura | 63.6 | Accuracy (%) | 86.4 |
| Pinocchio Italian - Diritto | 57.7 | Accuracy (%) | 86.4 |
| Open Italian LLM - MMLU-Pro (IT) | 29.15 | Accuracy (%) | 82.1 |
| Pinocchio Italian - Generale | 59.79 | Accuracy (%) | 79.5 |
| Pinocchio Italian - Matematica E Scienze | 49.38 | Accuracy (%) | 70.5 |
| Pinocchio Italian Leaderboard | 55.93 | Average Accuracy (%) | 70.5 |
| Pinocchio Italian - Logica | 42.84 | Accuracy (%) | 63.6 |
| EVALITA - admission-test | 55.62 | CPS | 55.3 |
| EVALITA - evalita NER | 33.53 | CPS | 51.1 |
| EVALITA - sentiment-analysis | 72.23 | CPS | 51.1 |
| EVALITA - text-entailment | 73.89 | CPS | 45.7 |
| EVALITA - summarization-fanpage | 25.69 | CPS | 42.6 |
Interactive version: theaggregate.ai/model?slug=maestrale-chat-v0-4-beta · How the rankings work · Data refreshed daily, snapshot 2026-07-22.