Clembench Multimodal v1.6.5 — leaderboard

Clembench Multimodal v1.6.5 evaluates model capability on agentic tasks from the linked upstream source with clemscore as the primary reported metric.

Metric: clemscore (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet80.77
2GPT-4o (2024-08-06)80.04
3GPT-473.55
4GPT-4o (2024-05-13)69.56
5Claude 3 Opus68.16
6Gemma 3 27B61.39
7GPT-4o Mini (2024-07-18)58.46
8Gemini 1.5 Flash47.73
9Pixtral-12B-240928.64
10InternVL2-8B23.17

Interactive version: theaggregate.ai/benchmark?slug=clembench-multimodal-v1-6-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.