Open Japanese LLM - Jcommonsenseqa Exact Match — leaderboard
Metric: Score (%). Source: huggingface.co. 418 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | aya-expanse-32B | 96.25 |
| 2 | EZO-Qwen2.5-32B-Instruct | 96.25 |
| 3 | Qwentile2.5-32B-Instruct | 96.07 |
| 4 | Qwen 2.5 32B | 95.98 |
| 5 | Qwen 2.5 14B | 95.98 |
| 6 | Awqward2.5-32B-Instruct | 95.89 |
| 7 | Qwen 2.5 32B Instruct | 95.8 |
| 8 | oxyge1-33B | 95.8 |
| 9 | Qwen2.5-32B-Instruct-CFT | 95.8 |
| 10 | Qwen2.5-32B-Instruct-abliterated-v2 | 95.71 |
| 11 | Qwen2.5-14B-Instruct-1M | 95.71 |
| 12 | lambda-qwen2.5-14B-dpo-test | 95.44 |
| 13 | DeepSeek R1 Distill Qwen 32B | 95.26 |
| 14 | tempmotacilla-cinerea-0308 | 95.26 |
| 15 | Qwen 2.5 14B Instruct | 95.17 |
Interactive version: theaggregate.ai/benchmark?slug=open-japanese-llm-jcommonsenseqa-exact-match · How the rankings work · Data refreshed daily, snapshot 2026-07-22.