J1-ENVS — leaderboard
Interactive legal-agent benchmark from J1Bench where agents complete Chinese legal consultation, drafting, civil court, and criminal court scenarios under procedural rules.
Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o (2024-11-20) | 63.9 |
| 2 | Qwen 3 32B | 59.14 |
| 3 | Gemma 3 27B | 55.73 |
| 4 | Llama 3.3 70B Instruct | 54.78 |
| 5 | DeepSeek V3 (0324) | 53.86 |
| 6 | glm-4-9B | 51.81 |
| 7 | Qwen 2.5 7B Instruct | 51.35 |
| 8 | Qwen 3 14B | 50.81 |
| 9 | Gemma 3 12B | 50.42 |
| 10 | Qwen 3 4B | 49.33 |
| 11 | Qwen 3 8B | 42.48 |
Interactive version: theaggregate.ai/benchmark?slug=j1-envs · How the rankings work · Data refreshed daily, snapshot 2026-07-22.