Yi Large: benchmark results
01.AI's closed API flagship chat model with strong multilingual performance, the Chinese startup's top proprietary offering of 2024. Provider: 01.AI. Released 2024-06-25. Access: API.
Unified ELO 1516 ± 1, rank #625 of 1392 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchBench | 78.89 | Aggregate Score (%) | 83.8 |
| MMLU | 79.3 | Accuracy (%) | 79.6 |
| WildBench | 48.93 | WB Score Task-Macro | 79 |
| BigCodeBench | 37.7 | Pass@1 (%) | 47.2 |
| ZebraLogic | 18.8 | Puzzle Accuracy (%) | 42.6 |
| BenchTable | 42.8 | Total Score (%) | 41 |
| AI for Education Pedagogy - Science | 74.86 | Accuracy (%) | 35.7 |
| AI for Education Pedagogy - Social studies | 70 | Accuracy (%) | 33.4 |
| TextClass Benchmark | 1473.21 | Meta-Elo (self-reported) | 31.8 |
| AI for Education Pedagogy | 71.52 | Accuracy (%) | 29.6 |
| AI for Education Pedagogy - Secondary | 70.6 | Accuracy (%) | 29.4 |
| AI for Education Pedagogy - Maths | 65.87 | Accuracy (%) | 25.8 |
Interactive version: theaggregate.ai/model?slug=yi-large · How It Works · Data refreshed daily, snapshot 2026-09-05.