Yi 6B (Chat) — benchmark results
Provider: 01.AI. Released 2023-11-02. Access: Open.
Unified ELO 1259 ± 23, rank #1601 of 1776 rated models, from 61 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM NaturalQuestions (Open) | 74.78 | F1 (%) | 73.3 |
| JustEval - Depth | 4.39 | Score (1-5) | 73.3 |
| SALAD-Bench Attack | 22.56 | Safety Score (%) | 72.7 |
| Open Chinese LLM - C-Eval Semantic | 75.31 | Accuracy (%) | 65.6 |
| Open Chinese LLM - CMMLU | 56.6 | Accuracy (%) | 62 |
| SALAD-Bench | 52.76 | Average Safety Score (%) | 51.5 |
| InfiBench | 38.14 | Score (%) | 51.4 |
| Open Chinese LLM - HellaSwag | 60.3 | Accuracy (%) | 50.7 |
| Open LLM Leaderboard - GPQA | 5.93 | Score | 49.7 |
| GSM8K | 44.9 | Accuracy (%) | 47.9 |
| Open Chinese LLM - ARC Challenge | 51.96 | Accuracy (%) | 47.8 |
| Open Chinese LLM - WinoGrande | 62.75 | Accuracy (%) | 47.8 |
Interactive version: theaggregate.ai/model?slug=yi-6b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.