MASLegalBench — leaderboard

Benchmark for evaluating multi-agent systems on deductive legal reasoning, with tasks grounded in GDPR-style rule application.

Metric: Best accuracy (self-reported). Source: benchmarklist.com. 5 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct86.21
2GPT-4o Mini84.32
3Qwen 2.5 7B Instruct75.47
4Qwen 3 8B68.74
5DeepSeek V3.163.05

Interactive version: theaggregate.ai/benchmark?slug=maslegalbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.