Human Expert — benchmark results
Expert-human baseline used as a reference point alongside model results. Provider: Human.
Unified ELO 2500 ± 48, rank #1 of 1776 rated models, from 72 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA AIME 2025 | 100 | Accuracy (%) | 100 |
| AA MATH-500 | 100 | Accuracy (%) | 100 |
| AA MMMU-Pro | 85.4 | Accuracy (%) | 100 |
| AISI Cyber Cooling Tower 100M | 7 | Avg Steps (/7) | 100 |
| AISI Cyber Cooling Tower 10M | 7 | Avg Steps (/7) | 100 |
| AISI Cyber TLO 100M | 32 | Avg Steps (/32) | 100 |
| AISI Cyber TLO 10M | 32 | Avg Steps (/32) | 100 |
| ARC Challenge (AI2) | 100 | Accuracy (%) | 100 |
| BALROG BabaIsAI (VLM) | 90 | Progress (%) | 100 |
| BALROG BabyAI (VLM) | 98 | Progress (%) | 100 |
| BALROG Crafter (VLM) | 90 | Progress (%) | 100 |
| BALROG MiniHack (VLM) | 80 | Progress (%) | 100 |
Interactive version: theaggregate.ai/model?slug=human-expert · How the rankings work · Data refreshed daily, snapshot 2026-07-22.