Volare — benchmark results
Moxoff's Italian-language SFT/LoRA fine-tune of Google's Gemma-7B, trained on SQuAD-it and in-house data for RAG and context-heavy tasks. Provider: Other. Released 2024-04-15. Access: Open.
Unified ELO 1415 ± 12, rank #1154 of 1776 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Pinocchio Italian - Lingua Straniera | 72.56 | Accuracy (%) | 79.5 |
| Pinocchio Italian - Matematica E Scienze | 50.96 | Accuracy (%) | 75 |
| Pinocchio Italian Leaderboard | 55.14 | Average Accuracy (%) | 61.4 |
| Pinocchio Italian - Cultura | 58.69 | Accuracy (%) | 59.1 |
| Pinocchio Italian - Generale | 55.83 | Accuracy (%) | 59.1 |
| Pinocchio Italian - Logica | 41.83 | Accuracy (%) | 59.1 |
| Pinocchio Italian - Diritto | 50.97 | Accuracy (%) | 54.5 |
| Open Italian LLM - MMLU-Pro (IT) | 23.61 | Accuracy (%) | 53.6 |
| EVALITA - lexical-substitution | 28 | CPS | 51.1 |
| EVALITA - evalita NER | 31.55 | CPS | 38.3 |
| EVALITA - summarization-fanpage | 23.28 | CPS | 29.8 |
| EVALITA - sentiment-analysis | 68.39 | CPS | 27.7 |
Interactive version: theaggregate.ai/model?slug=volare · How the rankings work · Data refreshed daily, snapshot 2026-07-22.