Median Human: benchmark results
Median performance for the measured general-human population, used as a reference point alongside model results. Provider: Human.
Unified ELO 1709 ± 17, rank #46 of 1392 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BabyVision | 94.1 | Score (%) | 100 |
| Epoch AI - Common Sense Qa 2 | 94.1 | Score | 100 |
| Epoch AI - Superglue | 89.8 | Score | 100 |
| GSM8K | 95 | Accuracy (%) | 100 |
| GameWorld Computer-Use | 64.1 | Progress (%) | 100 |
| GameWorld Generalist | 64.1 | Progress (%) | 100 |
| HellaSwag | 95.6 | Accuracy (%) | 100 |
| OpenBookQA | 92 | Accuracy (%) | 100 |
| PIQA | 95 | Accuracy (%) | 100 |
| WebArena | 78.24 | Success Rate (%) | 100 |
| WinoGrande | 94 | Accuracy (%) | 100 |
| SimpleBench | 83.7 | Score (AVG@5) | 99 |
Interactive version: theaggregate.ai/model?slug=median-human · How It Works · Data refreshed daily, snapshot 2026-09-05.