GenomeQA (Binary Choice): leaderboard

Metric: Accuracy (%) on the two-option (chance 50) GenomeQA questions over raw DNA sequences, unweighted mean of the six task families, one fixed system prompt, thinking enabled where supported; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)66.27
2GPT-5.1 (Thinking)60.77
3Claude Sonnet 4.5 (Thinking)59.7
4Grok 4.1 Fast (Reasoning)59.3
5Llama 4 Maverick Instruct56.6
6Qwen 3 Max (Preview) (Thinking)56.53

Interactive version: theaggregate.ai/benchmark?slug=genomeqa-binary-choice · How It Works · Data refreshed daily, snapshot 2026-10-07.