SciEval (K-12) - Macro-F1: leaderboard
Metric: Macro-F1 (%) of the predicted criterion score (0 not addressed to 3 extensive) against the expert score on the held-out test set of SciEval, a K-12 science instructional-material evaluation benchmark: given an NGSS lesson as page-marked PDF text and one EQuIP rubric criterion, the model returns a 0-3 score and the supporting evidence as JSON, simplified prompt, temperature 0; failed predictions are left out; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 4B 2507 Instruct | 29.5 |
| 2 | Gemini 2.0 Flash | 26.54 |
| 3 | GPT-4o Mini | 21.7 |
| 4 | Llama 3.1 8B Instruct | 18.18 |
| 5 | Llama 3.2 3B Instruct | 17.82 |
Interactive version: theaggregate.ai/benchmark?slug=scieval-k-12-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.