EduArt: leaderboard

Metric: Macro-average accuracy (%; mean over seven question formats of their format-specific scores: exact match for single-answer multiple choice, option-set F1 for multiple-answer choice and error identification, statement accuracy for true or false, element accuracy for positioning and blank accuracy for closed and open completion; 871 Italian and English questions, answer-only condition, one run; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)82.8
2GPT-5.580
3Gemini 3.5 Flash68.9
4Claude Opus 4.667.3
5Claude Sonnet 4.664.8
6Qwen 3 VL 235B A22B64.5
7Gemini 3.1 Flash Lite (Preview)62.8
8Mistral Large 358.8
9GPT-5.4 Mini54.9
10Claude Haiku 4.553.1
11GPT-5.4 Nano40.1
12Pixtral Large29.1

Interactive version: theaggregate.ai/benchmark?slug=eduart · How It Works · Data refreshed daily, snapshot 2026-09-29.