llm-jp-4-8B (Thinking): benchmark results
Provider: Other. Released 2026-03-16. Access: Open.
Unified ELO 1557 ± 1, rank #917 of 3078 rated models, from 45 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Japanese LLM v2 - JCoLA In-Domain | 79.65 | Accuracy (%) | 100 |
| Open Japanese LLM v2 - JCoLA Out-of-Domain | 81.35 | Accuracy (%) | 100 |
| Open Japanese LLM v2 - JBLiMP | 83.99 | Accuracy (%) | 92.9 |
| Open Japanese LLM v2 - M-IFEval-Ja | 45.35 | Prompt-Level Strict Accuracy (%) | 92.9 |
| Swallow - English MT-Bench - Extraction | 81.8 | Judge Score (normalized, %) | 80.6 |
| FrameBench - Frame Identification - Japanese | 75 | Accuracy (%; Japanese FrameNet candidate frames) | 80 |
| Swallow - Japanese MT-Bench - Writing | 67 | Judge Score (normalized, %) | 75.4 |
| Swallow - Post-trained Japanese - M-IFEval-Ja | 67.7 | Instruction-Level Strict Accuracy (%) | 70.1 |
| Swallow - Japanese MT-Bench - Extraction | 73.4 | Judge Score (normalized, %) | 69.4 |
| Swallow - Japanese MT-Bench - Roleplay | 69 | Judge Score (normalized, %) | 69.4 |
| Open Japanese LLM v2 - JFinQA | 61.3 | Accuracy (%) | 66.7 |
| Swallow - English MT-Bench - Writing | 73.7 | Judge Score (normalized, %) | 66.4 |
Interactive version: theaggregate.ai/model?slug=llm-jp-4-8b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.