LIBRA - LongContextMultiQ — leaderboard

Metric: Dataset Total Score (%). Source: huggingface.co. 17 models tracked.

Top models

#ModelScore
1GPT-4o36.67
2Llama 3.1 8B7.92
3Llama 2 7B7.92
4GLM-4 9B Chat7.75
5Llama 3 8B7
6Mistral-7B-v0.35.25
7Mistral Nemo5.17
8T-lite-instruct-0.15.17
9Mistral 7B4.83
10Mistral-7B-v0.14.42
11LongChat-7B-v1.5-32k3.17
12ChatGLM2 6B1.17

Interactive version: theaggregate.ai/benchmark?slug=libra-longcontextmultiq · How the rankings work · Data refreshed daily, snapshot 2026-07-22.