MBA-Bench - Specificity: leaderboard

Metric: Specificity rating (1-4; clarity and concreteness of the idea, scored by an InternVL2.5-78B MLLM judge over five business ideas for each of three questions per test image on the MBA-Bench multimodal business-ideation test set). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)4
2GPT-5 Mini4
3Gemini 3.5 Flash4
4Gemini 3.6 Flash4
5GPT-53.99
6Claude Sonnet 4.63.88
7Qwen 2.5 VL 7B Instruct3.6
8GPT-4o3.51
9Qwen 2.5 VL 32B Instruct3.5

Interactive version: theaggregate.ai/benchmark?slug=mba-bench-specificity · How It Works · Data refreshed daily, snapshot 2026-09-26.