Median Expert: benchmark results
Median performance for the relevant expert population, used as a reference point alongside model results. Provider: Human.
Unified ELO 1973 ± 20, rank #2 of 1392 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BALROG Crafter (VLM) | 90 | Progress (%) | 100 |
| BigCodeBench | 97 | Pass@1 (%) | 100 |
| ForecastBench | 70.7 | Overall Score (higher is better) | 100 |
| LAB-Bench FigQA No Tools (Opus 4.6 System Card) | 77 | Accuracy (%) | 100 |
| MMLU | 90 | Accuracy (%) | 100 |
| MMSI-Bench | 97.2 | Accuracy (%) | 100 |
| Video-MME-v2 | 94.94 | Avg Accuracy w/o sub (%) | 100 |
| Visual Aesthetic Benchmark (VAB) | 77.7 | Overall Top-1 ap@1 (self-reported) | 100 |
| VisualPuzzles | 89.3 | Overall Accuracy (%) | 100 |
| AA MMMU-Pro | 85.4 | Accuracy (%) | 98.1 |
| CodeElo | 1200 | Elo Rating | 91.2 |
| GeoBench ACW | 4100 | Average Score | 85.1 |
Interactive version: theaggregate.ai/model?slug=median-expert · How It Works · Data refreshed daily, snapshot 2026-09-05.