INTELLECT-3: benchmark results
Prime Intellect's open 106B A12B reasoning MoE, RL-trained from GLM-4.5-Air. Provider: PrimeIntellect. Released 2025-12-01. Access: Open.
Unified ELO 1573 ± 1, rank #347 of 1392 rated models, from 107 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - MedHallu Easy | 70.67 | Score (%) | 87.1 |
| AA LiveCodeBench | 77.67 | Pass@1 (%) | 86.9 |
| AA AIME 2025 | 88 | Accuracy (%) | 86.5 |
| Medmarks - HEAD-QA v2 | 88.3 | Score (%) | 85.7 |
| Medmarks - MedQA | 90.52 | Score (%) | 84.3 |
| Medmarks - MedConceptsQA Easy | 97.95 | Score (%) | 82.9 |
| Medmarks - LongHealth Task 1 | 88.67 | Score (%) | 82.1 |
| Medmarks - CareQA EN | 90.65 | Score (%) | 81.4 |
| Medmarks - MMLU-Pro Health | 76.69 | Score (%) | 81.4 |
| Medmarks - MedConceptsQA Medium | 74.1 | Score (%) | 81.4 |
| Medmarks - MedXpertQA Reasoning | 31.4 | Score (%) | 81.4 |
| Medmarks - MedXpertQA Understanding | 30.5 | Score (%) | 81.4 |
Interactive version: theaggregate.ai/model?slug=intellect-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.