LIBRA - LibrusecMHQA — leaderboard

Metric: Dataset Total Score (%). Source: huggingface.co. 17 models tracked.

Top models

#ModelScore
1GPT-4o50
2T-lite-instruct-0.148.44
3Llama 3 8B46.09
4Llama 3.1 8B45.31
5GLM-4 9B Chat44.53
6Mistral-7B-v0.339.06
7Mistral-7B-v0.134.11
8Mistral 7B33.59
9Mistral Nemo29.95
10Llama 2 7B27.6
11LongChat-7B-v1.5-32k24.74
12ChatGLM2 6B6.77

Interactive version: theaggregate.ai/benchmark?slug=libra-librusecmhqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.