J1-ENVS — leaderboard

Interactive legal-agent benchmark from J1Bench where agents complete Chinese legal consultation, drafting, civil court, and criminal court scenarios under procedural rules.

Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 17 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)63.9
2Qwen 3 32B59.14
3Gemma 3 27B55.73
4Llama 3.3 70B Instruct54.78
5DeepSeek V3 (0324)53.86
6glm-4-9B51.81
7Qwen 2.5 7B Instruct51.35
8Qwen 3 14B50.81
9Gemma 3 12B50.42
10Qwen 3 4B49.33
11Qwen 3 8B42.48

Interactive version: theaggregate.ai/benchmark?slug=j1-envs · How the rankings work · Data refreshed daily, snapshot 2026-07-22.