FigQA — leaderboard
FigQA evaluates model capability on multimodal tasks from the linked upstream source with Score as the primary reported metric.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Mythos 5 | 90.7 |
| 2 | Claude Mythos Preview | 89.3 |
| 3 | Claude Opus 4.8 | 87.3 |
| 4 | Claude Opus 4.7 | 85.4 |
Interactive version: theaggregate.ai/benchmark?slug=figqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.