J1-ENVS: leaderboard

Interactive legal-agent benchmark from J1Bench where agents complete Chinese legal consultation, drafting, civil court, and criminal court scenarios under procedural rules.

Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturated. 17 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)63.9
2Qwen 3 32B59.14
3Gemma 3 27B55.73
4Llama 3.3 70B Instruct54.78
5DeepSeek V353.86
6glm-4-9B51.81
7Qwen 2.5 7B Instruct51.35
8Qwen 3 14B50.81
9Gemma 3 12B50.42
10Qwen 3 4B49.33
11DeepSeek R143.48
12Qwen 3 8B42.48
13Mistral 7B26.71

Interactive version: theaggregate.ai/benchmark?slug=j1-envs · How It Works · Data refreshed daily, snapshot 2026-09-05.