MEGA-Bench Task - Math Convexity Value Estimation: leaderboard

Metric: Task Score (%). Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro (Preview 03-25)68.3
2Claude 3.5 Sonnet (20241022)63.4
3GPT-4o58.7
4Claude 3.5 Sonnet (20240620)57.5
5Gemini 1.5 Flash (002)57
6Gemini 1.5 Pro (002)55.6
7InternVL3-78B50.7
8Gemma 3 12B (IT)49.9
9Gemma 3 27B (IT)49.7
10InternVL3-38B49
11Qwen 2 VL 72B46.8
12GPT-4o Mini45.7
13InternVL2.5-78B45.2
14InternVL2-8B45
15Llama 4 Scout Base40.7

Interactive version: theaggregate.ai/benchmark?slug=mega-bench-task-math-convexity-value-estimation · How It Works · Data refreshed daily, snapshot 2026-09-05.