ELBench: leaderboard
Metric: Overall score (%; unweighted mean of the four module scores: general capability, safety and trustworthiness, basic education and high-level cultivation, each the mean per-item score; zero-shot, temperature 0; open-ended items scored by a Qwen3.6 rubric judge, majority of 9 calls). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Flash | 83.7 |
| 2 | Gemini 3.5 Flash | 83.4 |
| 3 | Seed 2.0 Pro | 83.2 |
| 4 | GPT-5.4 | 83.1 |
| 5 | DeepSeek V4 Pro | 83.1 |
| 6 | Claude Opus 4.8 | 83.1 |
| 7 | GLM-5.1 | 81.7 |
Interactive version: theaggregate.ai/benchmark?slug=elbench · How It Works · Data refreshed daily, snapshot 2026-09-29.