VRIQ (Abstract) - Sequence Completion: leaderboard

Metric: Accuracy (%) on the 200 abstract-domain VRIQ sequence completion puzzles (four-option IQ-style multiple choice over geometric primitives, adapted from public civil-service and aptitude exams or newly designed, solved by two annotators); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 14 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 32B (Thinking)44.5#287 (Qwen 3 VL 32B)
2Qwen 2.5 VL 7B Instruct40.5#643
3GPT-4o28.5#333
4GPT-5.128#131
5Gemini 2.5 Pro27#145
6GPT-5 Mini24.5#176
7GPT-4o Mini23.5#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=vriq-abstract-sequence-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.