Clembench Multimodal v1.6.5 — leaderboard
Clembench Multimodal v1.6.5 evaluates model capability on agentic tasks from the linked upstream source with clemscore as the primary reported metric.
Metric: clemscore (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet | 80.77 |
| 2 | GPT-4o (2024-08-06) | 80.04 |
| 3 | GPT-4 | 73.55 |
| 4 | GPT-4o (2024-05-13) | 69.56 |
| 5 | Claude 3 Opus | 68.16 |
| 6 | Gemma 3 27B | 61.39 |
| 7 | GPT-4o Mini (2024-07-18) | 58.46 |
| 8 | Gemini 1.5 Flash | 47.73 |
| 9 | Pixtral-12B-2409 | 28.64 |
| 10 | InternVL2-8B | 23.17 |
Interactive version: theaggregate.ai/benchmark?slug=clembench-multimodal-v1-6-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.