ELBench - General Capability: leaderboard

Metric: Module score (%; mean per-item score on 894 items sampled from MMLU-Pro, C-Eval, IFEval, MATH-500 and AIME 2024-2026, scored by reference matching or deterministic rules; zero-shot, temperature 0; open-ended items scored by a Qwen3.6 rubric judge, majority of 9 calls). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash93.4
2Claude Opus 4.891.9
3DeepSeek V4 Flash88.5
4DeepSeek V4 Pro88.2
5Seed 2.0 Pro86.6
6GPT-5.486.4
7GLM-5.183.2

Interactive version: theaggregate.ai/benchmark?slug=elbench-general-capability · How It Works · Data refreshed daily, snapshot 2026-09-29.