LinAlg-Bench - 4x4 Matrices: leaderboard
Metric: Overall accuracy (%) on the 220 integer-entry 4x4 matrix problems (nine task types: trace, transpose, matrix-vector product, multiplication, matrix power, nullity and rank, 20 problems each, plus 50 determinants and 30 eigenvalue problems), zero-shot chain-of-thought prompt at temperature 0 with tools and code execution disabled, one run per problem, answers checked against SymPy ground truth; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O1 | 97.3 |
| 2 | Gemini 3.1 Pro (Preview) | 93.2 |
| 3 | GPT-5.2 | 90 |
| 4 | DeepSeek V3.2 | 88.2 |
| 5 | Qwen 3 235B A22B 2507 Instruct | 86.8 |
| 6 | Mistral Large 3 | 84.5 |
| 7 | Claude Sonnet 4.5 | 80.9 |
| 8 | Llama 3.3 70B Instruct | 73.6 |
| 9 | GPT-4o | 63.6 |
| 10 | Qwen 2.5 72B Instruct | 62.7 |
Interactive version: theaggregate.ai/benchmark?slug=linalg-bench-4x4-matrices · How It Works · Data refreshed daily, snapshot 2026-10-07.