EduArt: leaderboard
Metric: Macro-average accuracy (%; mean over seven question formats of their format-specific scores: exact match for single-answer multiple choice, option-set F1 for multiple-answer choice and error identification, statement accuracy for true or false, element accuracy for positioning and blank accuracy for closed and open completion; 871 Italian and English questions, answer-only condition, one run; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 82.8 |
| 2 | GPT-5.5 | 80 |
| 3 | Gemini 3.5 Flash | 68.9 |
| 4 | Claude Opus 4.6 | 67.3 |
| 5 | Claude Sonnet 4.6 | 64.8 |
| 6 | Qwen 3 VL 235B A22B | 64.5 |
| 7 | Gemini 3.1 Flash Lite (Preview) | 62.8 |
| 8 | Mistral Large 3 | 58.8 |
| 9 | GPT-5.4 Mini | 54.9 |
| 10 | Claude Haiku 4.5 | 53.1 |
| 11 | GPT-5.4 Nano | 40.1 |
| 12 | Pixtral Large | 29.1 |
Interactive version: theaggregate.ai/benchmark?slug=eduart · How It Works · Data refreshed daily, snapshot 2026-09-29.