GenomeQA (Multiple Choice): leaderboard

Metric: Accuracy (%) on the four-option (chance 25) GenomeQA questions over raw DNA sequences, unweighted mean of the six task families, one fixed system prompt, thinking enabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)60.87
2Claude Sonnet 4.5 (Thinking)52.8
3GPT-5.1 (Thinking)52.3
4Grok 4.1 Fast (Reasoning)50.67
5Qwen 3 Max (Preview) (Thinking)45.33
6Llama 4 Maverick Instruct41.53

Interactive version: theaggregate.ai/benchmark?slug=genomeqa-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-10-07.