llm-jp-4-8B (Thinking): benchmark results

Provider: Other. Released 2026-03-16. Access: Open.

Unified ELO 1557 ± 1, rank #917 of 3078 rated models, from 45 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Japanese LLM v2 - JCoLA In-Domain79.65Accuracy (%)100
Open Japanese LLM v2 - JCoLA Out-of-Domain81.35Accuracy (%)100
Open Japanese LLM v2 - JBLiMP83.99Accuracy (%)92.9
Open Japanese LLM v2 - M-IFEval-Ja45.35Prompt-Level Strict Accuracy (%)92.9
Swallow - English MT-Bench - Extraction81.8Judge Score (normalized, %)80.6
FrameBench - Frame Identification - Japanese75Accuracy (%; Japanese FrameNet candidate frames)80
Swallow - Japanese MT-Bench - Writing67Judge Score (normalized, %)75.4
Swallow - Post-trained Japanese - M-IFEval-Ja67.7Instruction-Level Strict Accuracy (%)70.1
Swallow - Japanese MT-Bench - Extraction73.4Judge Score (normalized, %)69.4
Swallow - Japanese MT-Bench - Roleplay69Judge Score (normalized, %)69.4
Open Japanese LLM v2 - JFinQA61.3Accuracy (%)66.7
Swallow - English MT-Bench - Writing73.7Judge Score (normalized, %)66.4

Interactive version: theaggregate.ai/model?slug=llm-jp-4-8b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.