NuclearQAv2 - Verbal: leaderboard

Metric: Accuracy (%) on 283 short-answer concept questions, judged semantically equivalent to the reference by an LLM judge; mean of three runs; answers requested without reasoning text. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.282.33
2GPT-5.481.04
3NVIDIA-Nemotron-3-Super-120B-A12B-FP879.27
4GPT-OSS-120B78.33
5GPT-4o74.79
6Nova Pro70.79
7Llama 3 70B Instruct67.49
8Mistral 7B Instruct51.59
9Llama 3 8B Instruct48.65

Interactive version: theaggregate.ai/benchmark?slug=nuclearqav2-verbal · How It Works · Data refreshed daily, snapshot 2026-09-29.