ruGPT-3.5: benchmark results
Provider: Other. Access: Open.
Unified ELO 1182 ± 39, rank #1470 of 1537 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MERA Original Paper | 20.8 | Total Score (%; mean of 17 problem-solving and exam tasks, M | 72.2 |
| MERA Original Paper - SimpleAr | 2.9 | Exact Match (%; 5-shot, greedy generation) | 72.2 |
| MERA Original Paper - ruMultiAr | 2.5 | Exact Match (%; 5-shot, greedy generation) | 72.2 |
| MERA Original Paper - ruModAr | 0.1 | Exact Match (%; greedy generation) | 58.3 |
| MERA - LCS | 13.2 | Accuracy (%) | 53.1 |
| MERA Original Paper - MathLogicQA | 25.8 | Accuracy (%; 5-shot, log-likelihood) | 47.2 |
| MERA - CheGeKa | 14.47 | F1 (%) | 47.1 |
| MERA Original Paper - ruMMLU | 24.6 | Accuracy (%; 5-shot, log-likelihood) | 38.9 |
| MERA Original Paper - ruWorldTree | 24.6 | Accuracy (%; 5-shot, log-likelihood) | 38.9 |
| MERA - ruCodeEval | 0.24 | pass@1 (%) | 23.2 |
| MERA - USE | 8.24 | Grade, normalized (%) | 22.2 |
| MERA - ruHumanEval | 1.04 | pass@1 (%) | 21.9 |
Interactive version: theaggregate.ai/model?slug=rugpt-3-5 · How It Works · Data refreshed daily, snapshot 2026-09-25.