Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

Galileo Tool Tasks - xLAM Tool Missing — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 19 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet (20250219)96
2Claude 3.5 Sonnet (20241022)92
3Gemini 2.0 Flash (001)89
4Claude 3.5 Haiku (20241022)76
5O1 (2024-12-17)73
6Gemini 1.5 Flash71
7Qwen 2.5 72B Instruct66
8Mistral Large 2 (Nov) Instruct (2411)65
9O3 Mini (2025-01-31)63
10GPT-4o (2024-11-20)63
11Mistral Small 362
12Llama 3.3 70B Instruct61
13Gemini 1.5 Pro57
14GPT-4o Mini54
15Ministral-8B-Instruct-241034

Interactive version: theaggregate.ai/benchmark?slug=galileo-tool-tasks-xlam-tool-missing · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.