KOFFVQA - Commonsense Reasoning — leaderboard
Metric: Score (%). Source: huggingface.co. 82 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 (2025-04-16) | 96.22 |
| 2 | GPT-4.5 (Preview) | 93.56 |
| 3 | Gemini 2.5 Flash (Preview 04-17) | 92.67 |
| 4 | O1 (2024-12-17) | 92.67 |
| 5 | Gemini 2.5 Flash (Preview 05-20) | 92 |
| 6 | GPT-4.1 | 91.56 |
| 7 | GPT-4.1 Mini | 89.56 |
| 8 | Gemini 2.5 Pro (Preview 05-06) | 89.33 |
| 9 | Gemini 2.5 Pro (03-25) | 89.11 |
| 10 | Claude 3.5 Sonnet (20241022) | 88.89 |
| 11 | GPT-4o (2024-08-06) | 87.56 |
| 12 | Gemini 2.0 Pro (Preview 02-05) | 87.11 |
| 13 | Claude 3.7 Sonnet (20250219) | 86.89 |
| 14 | GPT-4o (2024-11-20) | 85.56 |
| 15 | Gemma 3 12B (IT) | 84.44 |
Interactive version: theaggregate.ai/benchmark?slug=koffvqa-commonsense-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.