Yi-1.5-9B — benchmark results
01.AI's Apache-2.0 9B Yi-1.5 base model (May 2024) with improved coding, math, reasoning, and instruction-following over the original Yi. Provider: 01.AI. Released 2024-05-13. Access: Open.
Unified ELO 1420 ± 9, rank #1126 of 1776 rated models, from 270 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - GPQA | 17.23 | Score | 94.5 |
| EuroEval English NLU - SQuAD | 86.89 | Reading comprehension Score (%) | 92.6 |
| EuroEval English NLU - SST-5 | 69.2 | Sentiment classification Score (%) | 92.4 |
| EuroEval German NLU - Germanquad | 66.85 | Reading comprehension Score (%) | 92.2 |
| EuroEval French NLU - FQuAD | 74.01 | Reading comprehension Score (%) | 91.4 |
| EuroEval Swedish NLU - Swerec | 79.08 | Sentiment classification Score (%) | 90.5 |
| EuroEval Dutch NLU - SQuAD NL | 79.05 | Reading comprehension Score (%) | 89.8 |
| EuroEval Finnish NLU - Tydiqa FI | 72.47 | Reading comprehension Score (%) | 89.8 |
| EuroEval Spanish NLU - MLQA ES | 66.85 | Reading comprehension Score (%) | 89 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 75.9 | Reading comprehension Score (%) | 88.2 |
| EuroEval Dutch NLU - DBRD | 91.36 | Sentiment classification Score (%) | 85.7 |
| Open Medical LLM - MMLU College Medicine | 67.63 | Accuracy (%) | 85.5 |
Interactive version: theaggregate.ai/model?slug=yi-1-5-9b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.