J1-ENVS: leaderboard
Interactive legal-agent benchmark from J1Bench where agents complete Chinese legal consultation, drafting, civil court, and criminal court scenarios under procedural rules.
Metric: Overall Score (self-reported). Source: benchmarklist.com. Status: saturated. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o (2024-11-20) | 63.9 |
| 2 | Qwen 3 32B | 59.14 |
| 3 | Gemma 3 27B | 55.73 |
| 4 | Llama 3.3 70B Instruct | 54.78 |
| 5 | DeepSeek V3 | 53.86 |
| 6 | glm-4-9B | 51.81 |
| 7 | Qwen 2.5 7B Instruct | 51.35 |
| 8 | Qwen 3 14B | 50.81 |
| 9 | Gemma 3 12B | 50.42 |
| 10 | Qwen 3 4B | 49.33 |
| 11 | DeepSeek R1 | 43.48 |
| 12 | Qwen 3 8B | 42.48 |
| 13 | Mistral 7B | 26.71 |
Interactive version: theaggregate.ai/benchmark?slug=j1-envs · How It Works · Data refreshed daily, snapshot 2026-09-05.