SciEval (K-12) - Macro-F1: leaderboard

Metric: Macro-F1 (%) of the predicted criterion score (0 not addressed to 3 extensive) against the expert score on the held-out test set of SciEval, a K-12 science instructional-material evaluation benchmark: given an NGSS lesson as page-marked PDF text and one EQuIP rubric criterion, the model returns a 0-3 score and the supporting evidence as JSON, simplified prompt, temperature 0; failed predictions are left out; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1Qwen 3 4B 2507 Instruct29.5
2Gemini 2.0 Flash26.54
3GPT-4o Mini21.7
4Llama 3.1 8B Instruct18.18
5Llama 3.2 3B Instruct17.82

Interactive version: theaggregate.ai/benchmark?slug=scieval-k-12-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.