| 1 | Anthropic | Claude Mythos 5.1 (Private) | 2054 ± 28 | 21 |
| 2 | Google | Gemini 4 Argon (Private) | 2015 ± 18 | 65 |
| 3 | OpenAI | GPT-6 Astra | 2011 ± 5 | 75 |
| 4 | xAI | Grok 4.7 | 1894 ± 9 | 20 |
| 5 | Meta | Muse Spark 1.3 | 1893 ± 7 | 72 |
| 6 | Alibaba | Qwen 3.8 Max (0902) | 1881 ± 13 | 182 |
| 7 | StepFun | Step 5 Preview | 1880 ± 16 | 6 |
| 8 | Moonshot | Kimi K3 | 1875 ± 4 | 10 |
| 9 | DeepSeek | DeepSeek V4.1 Flash | 1860 ± 7 | 39 |
| 9 | Zhipu | GLM-5.3 | 1860 ± 6 | 33 |
| 11 | Fireworks AI | Ember-1 (Low coverage) | 1858 ± 32 | 1 |
| 12 | Xiaomi | MiMo-V2.6-Pro | 1852 ± 12 | 7 |
| 13 | Tencent | Hy4 preview | 1845 ± 11 | 8 |
| 14 | ByteDance | Seed 2.1 Pro | 1818 ± 9 | 20 |
| 15 | InclusionAI | Ling 3.1 Flash | 1803 ± 39 | 10 |
| 16 | Apodex | Apodex 1.1 | 1798 ± 28 | 1 |
| 17 | Perplexity | Sonar Deep Research (Low coverage) | 1797 ± 32 | 3 |
| 18 | Cursor | Composer 2.5 (Low coverage) | 1777 ± 21 | 1 |
| 19 | Thinking Machines | Inkling | 1763 ± 8 | 2 |
| 20 | MiniMax | MiniMax-M3 | 1754 ± 4 | 11 |
| 21 | Nex AGI | Nex N2 Pro | 1751 ± 16 | 2 |
| 22 | Upstage | Solar Open2 250B | 1733 ± 27 | 8 |
| 23 | Meituan | LongCat-Flash-Thinking-2601 | 1717 ± 21 | 5 |
| 24 | Motif | Motif 3 | 1712 ± 21 | 1 |
| 25 | Kuaishou | KAT Coder Pro V2 | 1705 ± 49 | 2 |
| 25 | Skywork | Skywork-Reward-Gemma-2-27B-v0.2 (Low coverage) | 1705 ± 31 | 1 |
| 27 | Baidu | Ernie 5.1 | 1702 ± 20 | 9 |
| 28 | Applied Innovation Center | AIC-1 | 1699 ± 34 | 1 |
| 29 | NVIDIA | NVIDIA-Nemotron-3-Super-120B-A12B-FP8 | 1694 ± 18 | 20 |
| 30 | Mistral | Mistral-Large-3-675B-Instruct-2512 (Low coverage) | 1687 ± 21 | 60 |
| 31 | Poolside | Laguna S 2.1 | 1663 ± 20 | 4 |
| 32 | Shanghai AI Lab | Intern-S1 (Low coverage) | 1659 ± 19 | 29 |
| 33 | SK Telecom | A.X-K2 | 1658 ± 20 | 4 |
| 34 | SenseTime | SenseNova V6 Reasoner (Low coverage) | 1655 ± 40 | 1 |
| 35 | Cohere | command-a-reasoning-08-2025 (Low coverage) | 1651 ± 27 | 24 |
| 36 | Arcee AI | Arcee Trinity Large | 1648 ± 17 | 9 |
| 36 | Nous Research | Hermes 4 405B | 1648 ± 19 | 17 |
| 38 | iFlytek | Spark X2 (Low coverage) | 1646 ± 28 | 3 |
| 39 | KAUST | Hala-9B | 1632 ± 34 | 1 |
| 40 | LG AI | K-EXAONE | 1631 ± 14 | 8 |
| 41 | Amazon | Nova 2.0 Pro Preview | 1629 ± 48 | 11 |
| 42 | T-Bank | T-pro-it-2.0 (Low coverage) | 1628 ± 34 | 1 |
| 43 | TIGER-Lab | Qwen2.5-32B-Instruct-CFT | 1626 ± 27 | 3 |
| 44 | Qihoo 360 | 360Zhinao2-O1.5 | 1624 ± 35 | 2 |
| 45 | FreedomIntelligence | AceGPT-v2-70B-Chat | 1621 ± 26 | 4 |
| 45 | OpenBMB | MiniCPM-o-4.5 | 1621 ± 11 | 3 |
| 47 | Inception (G42), Cerebras & MBZUAI | Jais-2-70B-Chat | 1616 ± 31 | 2 |
| 48 | Inception | Mercury 2.5 | 1614 ± 27 | 2 |
| 48 | National Institute of Informatics | llm-jp-4-32B-a3B | 1614 ± 20 | 8 |
| 50 | Sarvam AI | Sarvam 105B | 1611 ± 34 | 3 |
| 51 | 01.AI | Yi Lightning | 1610 ± 18 | 23 |
| 52 | IBM | Granite 4.2 30B | 1606 ± 34 | 23 |
| 53 | PrimeIntellect | INTELLECT-3 | 1605 ± 18 | 1 |
| 54 | Institute of Science Tokyo | GPT-OSS-Swallow-20B-RL-v0.1 | 1604 ± 23 | 3 |
| 55 | Allen AI | Llama-3.1-Tulu-3-70B | 1596 ± 21 | 26 |
| 56 | AXCXEPT | EZO-Qwen2.5-32B-Instruct | 1592 ± 26 | 1 |
| 57 | INSAIT | MamayLM-Gemma-3-12B-IT-v1.0 | 1583 ± 39 | 1 |
| 58 | ABEJA | ABEJA-Qwen2.5-32B-Japanese-v1.0 | 1582 ± 20 | 1 |
| 59 | rinna | qwq-bakeneko-32B | 1581 ± 22 | 2 |
| 60 | China Mobile | JT-MINI | 1579 ± 50 | 1 |
| 61 | AI21 Labs | jamba-large-1.7 | 1577 ± 13 | 6 |
| 62 | CASIA & Wuhan AI | Taichu-VLR-7B | 1575 ± 24 | 2 |
| 63 | Return Zero | ko-gemma-2-9B-it (Low coverage) | 1573 ± 19 | 1 |
| 64 | ELYZA | ELYZA-Shortcut-1.0-Qwen-32B | 1570 ± 19 | 1 |
| 64 | UCSC-VLAA | VLAA-Thinker-Qwen2.5VL-7B | 1570 ± 24 | 2 |
| 66 | Sber | GigaChat 2 Max | 1568 ± 30 | 9 |
| 66 | SpeakLeash | Bielik-11B-v2.5-Instruct | 1568 ± 20 | 7 |
| 68 | AI Singapore | Llama 3.1 70B Cpt Sea Lionv3 Instruct | 1567 ± 26 | 5 |
| 69 | QCRI | Fanar-1-9B-Instruct | 1565 ± 22 | 2 |
| 70 | China Telecom | TeleChat2 | 1562 ± 45 | 1 |
| 70 | Sea AI Lab | Sailor2-20B-Chat | 1562 ± 21 | 3 |
| 72 | Princeton NLP | gemma-2-9B-it-SimPO | 1561 ± 16 | 48 |
| 72 | Reka | Reka Flash 3 | 1561 ± 21 | 1 |
| 74 | Sahabat-AI | gemma2-9B-cpt-sahabatai-v1-instruct (Low coverage) | 1556 ± 21 | 2 |
| 75 | Abacus.AI | Smaug-72B-v0.1 (Low coverage) | 1555 ± 24 | 8 |
| 76 | Axolotl | romulus-mistral-nemo-12B-simpo (Low coverage) | 1552 ± 20 | 1 |
| 77 | LLM360 | K2 Think V2 | 1551 ± 27 | 1 |
| 78 | A*STAR | MERaLiON-2-10B | 1549 ± 17 | 1 |
| 79 | UCLA | Gemma-2-9B-It-SPPO-Iter2 (Low coverage) | 1548 ± 20 | 7 |
| 80 | CyberAgent | DeepSeek-R1-Distill-Qwen-32B-Japanese | 1547 ± 19 | 1 |
| 80 | SB Intuitions | sarashina2-70B (Low coverage) | 1547 ± 48 | 7 |
| 82 | Inception (G42) & Cerebras | jais-adapted-70B-chat | 1538 ± 24 | 4 |
| 83 | Naver | HyperCLOVA X SEED Think (32B) | 1537 ± 22 | 1 |
| 84 | Lapa | lapa-v0.1.2-instruct | 1535 ± 39 | 2 |
| 85 | Cognitive Computations | dolphin-2.9.1-llama-3-70B (Low coverage) | 1533 ± 31 | 13 |
| 86 | Microsoft | Phi-3.5-MoE-instruct | 1531 ± 16 | 21 |
| 87 | University of Ljubljana | GaMS3-12B-Instruct | 1527 ± 22 | 1 |
| 88 | TII | Falcon-H1R-7B | 1525 ± 39 | 10 |
| 89 | Jon Durbin | Bagel-Hermes-34B-Slerp (Low coverage) | 1523 ± 24 | 3 |
| 90 | Linkbricks Horizon-AI | Linkbricks-Horizon-AI-Korean-Superb-22B | 1521 ± 18 | 1 |
| 91 | BAAI | Gemma2-9B-IT-Simpo-Infinity-Preference (Low coverage) | 1515 ± 21 | 9 |
| 92 | Tenyx | Llama3-TenyxChat-70B (Low coverage) | 1514 ± 24 | 1 |
| 93 | PLLuM | Llama-PLLuM-70B-chat (Low coverage) | 1513 ± 23 | 6 |
| 94 | SDAIA | ALLaM-7B-Instruct-preview | 1507 ± 19 | 1 |
| 95 | omlab | VLM-R1-3B-Math-0305 | 1503 ± 24 | 1 |
| 96 | VAGOsolutions | Llama-3.1-SauerkrautLM-8B-Instruct (Low coverage) | 1501 ± 20 | 5 |
| 97 | KT | Midm-2.0-Base-Instruct | 1498 ± 30 | 2 |
| 98 | URSA-MATH | URSA-8B | 1494 ± 22 | 1 |
| 99 | Liquid AI | LFM2.5-VL-1.6B | 1491 ± 27 | 12 |
| 100 | Apertus | Apertus-70B-Instruct-2509 | 1486 ± 18 | 2 |
| 101 | Valiant Labs | Llama 3 70B Fireplace (Low coverage) | 1484 ± 24 | 4 |
| 102 | OrionStar | OrionStar-Yi-34B-Chat | 1482 ± 34 | 1 |
| 102 | SILMA AI | SILMA-9B-Instruct-v1.0 | 1482 ± 24 | 1 |
| 104 | Yanolja | EEVE-Korean-Instruct-10.8B-v1.0 (Low coverage) | 1479 ± 18 | 3 |
| 105 | HuggingFace | starchat2-15B-v0.1 | 1470 ± 21 | 14 |
| 106 | Stanford CRFM | Marin 8B Instruct (Low coverage) | 1467 ± 84 | 2 |
| 107 | Preferred Elements | plamo-2-8B (Low coverage) | 1466 ± 48 | 2 |
| 108 | Nanbeige | Nanbeige4.1-3B | 1465 ± 38 | 2 |
| 109 | Argilla | notux-8x7B-v1 (Low coverage) | 1461 ± 31 | 1 |
| 109 | Lightblue | suzume-llama-3-8B-multilingual-orpo-borda-top75 (Low coverage) | 1461 ± 18 | 5 |
| 111 | Kakao | kanana-2-30B-a3B | 1460 ± 30 | 1 |
| 112 | OpenBuddy | openbuddy-qwen2.5llamaify-14B-v23.1-200k (Low coverage) | 1459 ± 20 | 7 |
| 113 | Saltlux | luxia-21.4B-alignment-v1.0 (Low coverage) | 1457 ± 19 | 2 |
| 114 | RLHFlow | LLaMA3-iterative-DPO-final (Low coverage) | 1452 ± 19 | 2 |
| 115 | University of Bari | LLaMAntino-3-ANITA-8B-Inst-DPO-ITA (Low coverage) | 1450 ± 22 | 1 |
| 116 | OpenChat | openchat-3.5-1210 | 1447 ± 21 | 5 |
| 117 | Databricks | DBRX | 1446 ± 12 | 7 |
| 117 | Korea University | KULLM3 (Low coverage) | 1446 ± 19 | 1 |
| 119 | CMU | MAmmoTH2-8B-Plus | 1441 ± 15 | 3 |
| 120 | Bllossom | llama-3-Korean-Bllossom-8B (Low coverage) | 1436 ± 18 | 1 |
| 120 | MTS AI | multi_verse_model (Low coverage) | 1436 ± 19 | 1 |
| 122 | Stability AI | StableBeluga2 | 1434 ± 11 | 13 |
| 123 | Nexusflow | Starling-LM-7B-beta | 1430 ± 12 | 2 |
| 124 | Salesforce | Llama 3 8B SFR Iterative DPO R (Low coverage) | 1428 ± 18 | 2 |
| 125 | Magpie | Llama 3 8B Magpie Align V0.3 (Low coverage) | 1423 ± 18 | 5 |
| 126 | UC Berkeley | Starling-LM-7B-alpha | 1411 ± 12 | 1 |
| 127 | SCB 10X | llama-3-typhoon-v1.5-8B-instruct | 1410 ± 20 | 1 |
| 128 | Rakuten | RakutenAI-7B-chat | 1409 ± 30 | 1 |
| 129 | Tsinghua & ByteDance | SALMONN-13B (Low coverage) | 1407 ± 21 | 2 |
| 130 | Moxoff | Volare (Low coverage) | 1405 ± 23 | 1 |
| 131 | Intel | neural-chat-7B-v3-3 (Low coverage) | 1399 ± 26 | 2 |
| 132 | Teknium | OpenHermes-2.5-Mistral-7B | 1398 ± 18 | 3 |
| 133 | Deci AI | DeciLM-7B-instruct (Low coverage) | 1394 ± 29 | 2 |
| 134 | Infly | OpenCoder-8B-Instruct (Low coverage) | 1393 ± 24 | 1 |
| 135 | Baichuan | Baichuan2-13B-Chat | 1391 ± 16 | 3 |
| 136 | GritLM | GritLM-7B (Low coverage) | 1386 ± 18 | 2 |
| 136 | KAIST | janus-dpo-7B (Low coverage) | 1386 ± 30 | 4 |
| 138 | Tilde | TildeOpen-30B | 1382 ± 30 | 1 |
| 139 | Writer | Palmyra X (43B) (Low coverage) | 1369 ± 24 | 1 |
| 140 | OpenLLM-France | Lucie-7B | 1367 ± 29 | 1 |
| 141 | LMSYS | vicuna-33B-v1.3 | 1363 ± 10 | 9 |
| 142 | Open-Orca | Mistral-7B-OpenOrca (Low coverage) | 1360 ± 16 | 1 |
| 143 | VoiceLab | trurl-2-13B-academic | 1339 ± 11 | 2 |
| 144 | Kurakura AI | Luth-0.6B-Instruct | 1332 ± 30 | 1 |
| 145 | NTQ | Nxcode-CQ-7B-orpo (Low coverage) | 1326 ± 20 | 1 |
| 146 | Mobius Labs | DeepSeek-R1-ReDistill-Llama3-8B-v1.1 | 1311 ± 26 | 2 |
| 147 | H2O.ai | h2o-danube3-4B-chat (Low coverage) | 1308 ± 57 | 4 |
| 148 | Johns Hopkins University | ALMA-13B-Pretrain (Low coverage) | 1306 ± 26 | 2 |
| 149 | Avignon Univ | BioMistral-7B | 1298 ± 16 | 1 |
| 150 | BigCode | starcoder2-15B | 1291 ± 18 | 4 |
| 151 | Together | Llama 2 7B 32K | 1266 ± 12 | 6 |
| 152 | RWKV | RWKV-6-World-7B (Low coverage) | 1261 ± 45 | 1 |
| 153 | Sakana AI | TinySwallow-1.5B (Low coverage) | 1238 ± 33 | 1 |
| 154 | AI Sweden | gpt-sw3-40B | 1231 ± 29 | 4 |
| 155 | Singapore Univ | TinyLlama-1.1B-Chat-v1.0 | 1226 ± 37 | 2 |
| 156 | Phind | Phind-CodeLlama-34B-v2 | 1195 ± 14 | 1 |
| 157 | EleutherAI | gpt-neox-20B | 1186 ± 10 | 13 |
| 158 | BigScience | BLOOMZ-7b1 (Low coverage) | 1180 ± 27 | 5 |
| 159 | NucleusAI | nucleus-22B-token-500B (Low coverage) | 1179 ± 33 | 1 |
| 160 | State Space Models | mamba-2.8B-hf | 1095 ± 45 | 1 |
| 161 | CroissantLLM | CroissantLLMChat-v0.1 | 1030 ± 31 | 1 |
| Unranked | AI4Bharat | No ranked model | — | 0 |
| Unranked | Boson AI | No ranked model | — | 0 |
| Unranked | BSC | No ranked model | — | 0 |
| Unranked | Celeris | No ranked model | — | 0 |
| Unranked | Cerebras | No ranked model | — | 0 |
| Unranked | Cisco | No ranked model | — | 0 |
| Unranked | CUHK Shenzhen | No ranked model | — | 0 |
| Unranked | DataCanvas | No ranked model | — | 0 |
| Unranked | DeepAuto | No ranked model | — | 0 |
| Unranked | Du Xiaoman | No ranked model | — | 0 |
| Unranked | EPFL | No ranked model | — | 0 |
| Unranked | EuroLLM | No ranked model | — | 0 |
| Unranked | FuseAI | No ranked model | — | 0 |
| Unranked | HUST | No ranked model | — | 0 |
| Unranked | John Snow Labs | No ranked model | — | 0 |
| Unranked | Krutrim | No ranked model | — | 0 |
| Unranked | LAION | No ranked model | — | 0 |
| Unranked | Langboat | No ranked model | — | 0 |
| Unranked | LLaVA Team | No ranked model | — | 0 |
| Unranked | LM Provers | No ranked model | — | 0 |
| Unranked | Manus | No ranked model | — | 0 |
| Unranked | MedAlpaca | No ranked model | — | 0 |
| Unranked | mii-llm | No ranked model | — | 0 |
| Unranked | Numina & Hugging Face | No ranked model | — | 0 |
| Unranked | NYU | No ranked model | — | 0 |
| Unranked | Occiglot | No ranked model | — | 0 |
| Unranked | OpenGPT-X | No ranked model | — | 0 |
| Unranked | OpenLM Research | No ranked model | — | 0 |
| Unranked | OpenThaiGPT | No ranked model | — | 0 |
| Unranked | OpenThoughts | No ranked model | — | 0 |
| Unranked | Ornith | No ranked model | — | 0 |
| Unranked | RedNote | No ranked model | — | 0 |
| Unranked | Rhymes AI | No ranked model | — | 0 |
| Unranked | Saama AI Labs | No ranked model | — | 0 |
| Unranked | Samsung | No ranked model | — | 0 |
| Unranked | SenseLLM | No ranked model | — | 0 |
| Unranked | SentientAGI | No ranked model | — | 0 |
| Unranked | ServiceNow | No ranked model | — | 0 |
| Unranked | TensorOpera | No ranked model | — | 0 |
| Unranked | The Fin AI | No ranked model | — | 0 |
| Unranked | Trillion Labs | No ranked model | — | 0 |
| Unranked | UCAS | No ranked model | — | 0 |
| Unranked | UCSD | No ranked model | — | 0 |
| Unranked | UIUC | No ranked model | — | 0 |
| Unranked | USTC | No ranked model | — | 0 |
| Unranked | VAIV | No ranked model | — | 0 |
| Unranked | Various | No ranked model | — | 0 |
| Unranked | Vercel | No ranked model | — | 0 |
| Unranked | Vikhr | No ranked model | — | 0 |
| Unranked | vikhyatk | No ranked model | — | 0 |
| Unranked | VinAI | No ranked model | — | 0 |
| Unranked | vivo | No ranked model | — | 0 |
| Unranked | WeChat AI | No ranked model | — | 0 |
| Unranked | WinGPT | No ranked model | — | 0 |
| Unranked | XVERSE | No ranked model | — | 0 |
| Unranked | Yandex | No ranked model | — | 0 |
| Unranked | Zhejiang University | No ranked model | — | 0 |
|---|