Benchmark Catalog
5615 public benchmarks tracked with scores, source links, and saturation status. The most broadly covered benchmarks are listed below.
| Benchmark | Category | Metric | Models | Saturation |
|---|---|---|---|---|
| Open LLM Leaderboard - BBH | Score | 586 | ||
| Open LLM Leaderboard - GPQA | Score | 586 | ||
| Open LLM Leaderboard - IFEval | Score | 586 | ||
| Open LLM Leaderboard - MATH Level 5 | Score | 586 | ||
| Open LLM Leaderboard - MMLU-Pro | Score | 586 | ||
| Open LLM Leaderboard - MuSR | Score | 586 | ||
| Artificial Analysis Intelligence Index | Intelligence Index | 482 | ||
| AA GPQA Diamond | Math, Logic & Reasoning | Accuracy (%) | 477 | saturated |
| AA Humanity's Last Exam | General Knowledge & Language | Accuracy (%) | 468 | years away from saturation |
| AA CritPt | Math, Logic & Reasoning | Accuracy (%) | 413 | saturation imminent |
| AA Omniscience | General Knowledge & Language | Score | 408 | saturation imminent |
| AA Omniscience - Business | Accuracy (%) | 408 | ||
| AA Omniscience - Health | Accuracy (%) | 408 | ||
| AA Omniscience - Humanities & Social Sciences | Accuracy (%) | 408 | ||
| AA Omniscience - Law | Accuracy (%) | 408 | ||
| AA Omniscience - Software Engineering (SWE) | Accuracy (%) | 408 | ||
| AA-Omniscience Accuracy | Accuracy (%) | 408 | ||
| AA Omniscience - Science, Engineering & Mathematics | Accuracy (%) | 407 | ||
| AA Long Context Reasoning | General Knowledge & Language | Accuracy (%) | 404 | saturation imminent |
| BenchmarkList ECI | Capability Index (ECI) | 364 | ||
| AA IFBench | General Knowledge & Language | Accuracy (%) | 345 | years away from saturation |
| AA TAU-2 Bench | Agentic & Tool Use | Accuracy (%) | 337 | years away from saturation |
| AA Terminal-Bench Hard | Coding & Software Engineering | Accuracy (%) | 331 | saturation imminent |
| UGI - Natural Intelligence | NatInt Score | 315 | ||
| UGI - Willingness (W/10) | W/10 Score | 315 | ||
| UGI Leaderboard | UGI Score | 315 | ||
| UGI - Writing | Writing Score | 307 | ||
| Chatbot Arena (Text - English) | Arena Score | 304 | ||
| Chatbot Arena (Text - Hard Prompts) | Arena Score | 304 | ||
| Chatbot Arena (Text - Instruction Following) | Arena Score | 304 | ||
| Chatbot Arena (Text) | Elo | 304 | ||
| LLM Stats Score | LLM Stats Score (conservative rating) | 304 | ||
| Chatbot Arena (Text - Coding) | Arena Score | 303 | ||
| Chatbot Arena (Text - Creative Writing) | Arena Score | 303 | ||
| Chatbot Arena (Text - Multi-Turn) | Arena Score | 303 | ||
| Wolfram LLM Benchmarking Project | Coding & Software Engineering | Correct Functionality (%) | 297 | years away from saturation |
| Chatbot Arena (Text - Longer Query) | Arena Score | 295 | ||
| Chatbot Arena (Text - Math) | Arena Score | 295 | ||
| Chatbot Arena (Text - Chinese) | Arena Score | 289 | ||
| Chatbot Arena (Text - Russian) | Arena Score | 289 | ||
| SpeechMap Compliance | % Requests Completed | 286 | ||
| AI Chess Leaderboard (Reasoning) | Games & Puzzles | Elo | 285 | saturation imminent |
| Open Chinese LLM - ARC Challenge | Accuracy (%) | 258 | ||
| Open Chinese LLM - C-Eval Semantic | Accuracy (%) | 258 | ||
| Open Chinese LLM - CMMLU | Accuracy (%) | 258 | ||
| Open Chinese LLM - GSM8K | Accuracy (%) | 258 | ||
| Open Chinese LLM - HellaSwag | Accuracy (%) | 258 | ||
| Open Chinese LLM - TruthfulQA MC | Accuracy (%) | 258 | ||
| Open Chinese LLM - WinoGrande | Accuracy (%) | 258 | ||
| Open Chinese LLM Leaderboard | General Knowledge & Language | Average Score (%) | 258 | saturated |
| OTIS Mock AIME 2024-25 | Math, Logic & Reasoning | Accuracy (%) | 248 | saturated |
| Chatbot Arena (Text - German) | Arena Score | 245 | ||
| CritPt | Science & Domain Knowledge | Accuracy (self-reported) | 242 | saturation imminent |
| AI Chess Leaderboard (Continuation) | Games & Puzzles | Elo | 240 | years away from saturation |
| LM Market Cap LMC Score | LMC Score (0-100) | 237 | ||
| Chatbot Arena (Text - French) | Arena Score | 236 | ||
| Chatbot Arena (Text - Spanish) | Arena Score | 236 | ||
| EuroEval Dutch NLU - DBRD | Sentiment classification Score (%) | 232 | ||
| EuroEval Icelandic Knowledge | Knowledge Average Score (%) | 228 | ||
| EuroEval Norwegian Common Sense Reasoning | Common Sense Reasoning Average Score (%) | 228 |
Interactive version: theaggregate.ai/benchmarks · How It Works · Data refreshed daily, snapshot 2026-09-05.