Jais-2-70B-Chat: benchmark results
Provider: Other. Released 2025-12-09. Access: Open.
Unified ELO 1692 ± 36, rank #348 of 2656 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| QIMMA - ArabCulture | 83.24 | Accuracy (%, log-likelihood multiple choice) | 100 |
| QIMMA - ArabicMMLU | 81.29 | Accuracy (%, log-likelihood multiple choice) | 100 |
| QIMMA - AraTrust | 90.23 | Accuracy (%, log-likelihood multiple choice) | 98.8 |
| QIMMA - MizanQA | 71.78 | Normalized multiple-choice probability score (%) | 98.8 |
| QIMMA - PALMX | 83.73 | Accuracy (%, log-likelihood multiple choice) | 98.8 |
| QIMMA - AraDiCE-Culture | 78.89 | Accuracy (%, log-likelihood multiple choice) | 97 |
| QIMMA - MedAraBench | 50.89 | Accuracy (%, log-likelihood multiple choice) | 95.1 |
| QIMMA - Overall | 67.77 | Mean score across 14 benchmarks (%) | 95.1 |
| QIMMA - MedArabiQ QA | 54.67 | BERTScore F1 (%) | 93.9 |
| QIMMA - 3LM STEM | 87.96 | Accuracy (%, log-likelihood multiple choice) | 89 |
| QIMMA - MedArabiQ MCQ | 50.91 | Accuracy (%, log-likelihood multiple choice) | 89 |
| QIMMA - GAT | 51.67 | Accuracy (%, log-likelihood multiple choice) | 84.1 |
Interactive version: theaggregate.ai/model?slug=jais-2-70b-chat · How It Works · Data refreshed daily, snapshot 2026-09-19.