MERA Original Paper: leaderboard

Metric: Total Score (%; mean of 17 problem-solving and exam tasks, MERA code v1.1.0). Source: arxiv.org. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScore
1Mistral-7B-v0.140
2davinci-00238.3
3Llama 2 13B Base36.8
4Yi 6B (Base)35.4
5Llama 2 7B Base32.7
6ruGPT-3.520.8
7ruGPT-3 (Medium)20.1
8mGPT19.8
9mGPT-13B19.6
10ruGPT3-large19.3
11ruT5-base19.3
12ruGPT-3-small19.1

Interactive version: theaggregate.ai/benchmark?slug=mera-original-paper · How It Works · Data refreshed daily, snapshot 2026-09-25.