ELBench: leaderboard

Metric: Overall score (%; unweighted mean of the four module scores: general capability, safety and trustworthiness, basic education and high-level cultivation, each the mean per-item score; zero-shot, temperature 0; open-ended items scored by a Qwen3.6 rubric judge, majority of 9 calls). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1DeepSeek V4 Flash83.7
2Gemini 3.5 Flash83.4
3Seed 2.0 Pro83.2
4GPT-5.483.1
5DeepSeek V4 Pro83.1
6Claude Opus 4.883.1
7GLM-5.181.7

Interactive version: theaggregate.ai/benchmark?slug=elbench · How It Works · Data refreshed daily, snapshot 2026-09-29.