FoundationalASSIST - Difficulty Comparison: leaderboard

Metric: Accuracy (%; which of two problems is harder by IRT difficulty from the full response data, about 10,000 stratified pairs; chance 50). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1GPT-OSS-120B68.6
2Llama 3.3 70B Instruct65.7
3Qwen 3 Next 80B A3B Instruct63.9
4Qwen 3 Next 80B A3B (Thinking)63.7

Interactive version: theaggregate.ai/benchmark?slug=foundationalassist-difficulty-comparison · How It Works · Data refreshed daily, snapshot 2026-09-26.