youri-7B-chat: benchmark results
Provider: Other. Access: Open.
Unified ELO 1353 ± 21, rank #2466 of 2928 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - QA Mode - Truthfulness | 0.73 | Truthfulness Score | 46.7 |
| Open LLM Leaderboard v1 - WinoGrande | 75.06 | Accuracy (%) (5-shot) | 40.1 |
| pfgen-bench - QA Mode - Score | 0.44 | pfgen Score (mean of three) | 37 |
| pfgen-bench - QA Mode - Fluency | 0.5 | Fluency Score | 32.1 |
| pfgen-bench - QA Mode - Helpfulness | 0.08 | Helpfulness Score | 29.7 |
| Open LLM Leaderboard v1 - HellaSwag | 76.09 | Normalized accuracy (%) (10-shot) | 28 |
| Open LLM Leaderboard v1 - ARC Challenge | 51.19 | Normalized accuracy (%) (25-shot) | 25.6 |
| Open LLM Leaderboard v1 - MMLU | 46.06 | Accuracy (%) (5-shot) | 24.6 |
| Open LLM Leaderboard v1 - GSM8K | 1.52 | Accuracy (%) (5-shot) | 17.3 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 41.17 | MC2 (%) (0-shot) | 16.3 |
Interactive version: theaggregate.ai/model?slug=youri-7b-chat · How It Works · Data refreshed daily, snapshot 2026-09-23.