Structured Local-Deployment MCQ: leaderboard
Metric: Strict accuracy (%): share of the 1,085 items answered with the correct option letter in a format-valid output (a correct but malformed answer scores zero), deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 1.5B Instruct | 67.1 |
| 2 | Qwen 3.5 2B | 64.98 |
| 3 | stablelm-zephyr-3B | 54.19 |
| 4 | SmolLM2-1.7B-Instruct | 42.76 |
| 5 | SmolLM2-360M-Instruct | 32.07 |
Interactive version: theaggregate.ai/benchmark?slug=structured-local-deployment-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.