MERA Original Paper: leaderboard
Metric: Total Score (%; mean of 17 problem-solving and exam tasks, MERA code v1.1.0). Source: arxiv.org. Saturation forecast: Estimated already saturated. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Mistral-7B-v0.1 | 40 |
| 2 | davinci-002 | 38.3 |
| 3 | Llama 2 13B Base | 36.8 |
| 4 | Yi 6B (Base) | 35.4 |
| 5 | Llama 2 7B Base | 32.7 |
| 6 | ruGPT-3.5 | 20.8 |
| 7 | ruGPT-3 (Medium) | 20.1 |
| 8 | mGPT | 19.8 |
| 9 | mGPT-13B | 19.6 |
| 10 | ruGPT3-large | 19.3 |
| 11 | ruT5-base | 19.3 |
| 12 | ruGPT-3-small | 19.1 |
Interactive version: theaggregate.ai/benchmark?slug=mera-original-paper · How It Works · Data refreshed daily, snapshot 2026-09-25.