LIBRA - ru2WikiMultihopQA — leaderboard

Metric: Dataset Total Score (%). Source: huggingface.co. 17 models tracked.

Top models

#ModelScore
1GPT-4o76.67
2GLM-4 9B Chat48.78
3Mistral 7B43.21
4Mistral-7B-v0.341
5Llama 2 7B37.19
6LongChat-7B-v1.5-32k35.16
7Llama 3.1 8B33.41
8Mistral Nemo27.94
9Mistral-7B-v0.122.99
10Llama 3 8B18.37
11ChatGLM2 6B17.48
12T-lite-instruct-0.112.93

Interactive version: theaggregate.ai/benchmark?slug=libra-ru2wikimultihopqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.