Midm-2.0-Base-Instruct: benchmark results
Provider: Other.
Unified ELO 1498 ± 31, rank #765 of 1607 rated models, from 39 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Polar (US) - Sociocultural Axis | 84.25 | Sociocultural-axis ICAT (0-100), the StereoSet idealized con | 75.7 |
| Horangi 4 - KorQuAD 1.0 | 85.15 | F1 (x100) | 73.1 |
| Horangi 4 - KoBBQ | 92 | Accuracy (%) | 72.6 |
| Polar (South Korea) - Economic Axis | 87.24 | Economic-axis ICAT (0-100), the StereoSet idealized context | 67.6 |
| Polar (US) | 76.12 | Total ICAT (0-100), the StereoSet idealized context associat | 64.9 |
| Polar (South Korea) | 88.17 | Total ICAT (0-100), the StereoSet idealized context associat | 59.5 |
| Polar (South Korea) - Sociocultural Axis | 89.09 | Sociocultural-axis ICAT (0-100), the StereoSet idealized con | 51.4 |
| Horangi 4 - HAE-RAE Bench (without Reading Comprehension) | 81.82 | Accuracy (%) | 34.6 |
| Polar (US) - Economic Axis | 67.99 | Economic-axis ICAT (0-100), the StereoSet idealized context | 32.4 |
| Horangi 4 - Ko-HalluLens (Nonexistent Entities) | 64 | Refusal rate (%) | 32.2 |
| Horangi 4 - Ko-HLE | 10 | Accuracy (%) | 29.3 |
| Horangi 4 - GLP - General Knowledge | 73.91 | Score (%) | 28.8 |
Interactive version: theaggregate.ai/model?slug=midm-2-0-base-instruct · How It Works · Data refreshed daily, snapshot 2026-09-29.