vicuna-13B-v1.5-16k: benchmark results
Provider: LMSYS. Access: Open.
Unified ELO 1400 ± 20, rank #2144 of 2928 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 51.96 | MC2 (%) (0-shot) | 53.9 |
| pfgen-bench - QA Mode - Fluency | 0.57 | Fluency Score | 49.7 |
| pfgen-bench - QA Mode - Score | 0.46 | pfgen Score (mean of three) | 41.2 |
| Open LLM Leaderboard v1 - HellaSwag | 80.37 | Normalized accuracy (%) (10-shot) | 40.8 |
| pfgen-bench - QA Mode - Truthfulness | 0.72 | Truthfulness Score | 40.6 |
| Open LLM Leaderboard v1 - ARC Challenge | 56.74 | Normalized accuracy (%) (25-shot) | 36.1 |
| Open LLM Leaderboard v1 - MMLU | 55.28 | Accuracy (%) (5-shot) | 35.9 |
| pfgen-bench - QA Mode - Helpfulness | 0.1 | Helpfulness Score | 35.8 |
| Open LLM Leaderboard v1 - GSM8K | 13.12 | Accuracy (%) (5-shot) | 35 |
| Open LLM Leaderboard v1 - WinoGrande | 72.38 | Accuracy (%) (5-shot) | 29 |
| SALAD-Bench | 24.78 | Average Safety Score (%) | 3 |
| SALAD-Bench Attack | 3.66 | Safety Score (%) | 3 |
Interactive version: theaggregate.ai/model?slug=vicuna-13b-v1-5-16k · How It Works · Data refreshed daily, snapshot 2026-09-23.