llm leaderboard

Best LLM for math

Research-grade problem sets plus competition math. Contamination-resistant private sets carry the same weight as public ones.

Aldena runs these models inside your team rooms. See what each one costs.

378 of 378 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarks%%%%%elo
1
Claude Fable 5 (max)anthropic/claude-fable-5:max
Anthropic82.0$10.00 / $50.004 of 6 benchmarks100.0%87.0%100.0%99.7%
2
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic80.5$5.00 / $25.004 of 6 benchmarks85.6%73.2%98.9%1556
3
GPT-5.6 Sol (max)openai/gpt-5.6-sol:max
OpenAI79.3$5.00 / $30.003 of 6 benchmarks89.1%82.9%100.0%
4
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI78.8$2.50 / $15.004 of 6 benchmarks80.0%50.0%97.8%1494
5
GPT-5.6 Terra (max)openai/gpt-5.6-terra:max
OpenAI77.0$1.00 / $6.003 of 6 benchmarks86.0%70.7%99.7%
6
Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max
Anthropic74.8$5.00 / $25.004 of 6 benchmarks47.2%80.0%56.1%98.3%
7
GPT-5.6 Luna (max)openai/gpt-5.6-luna:max
OpenAI74.6$0.10 / $0.603 of 6 benchmarks82.1%61.0%98.3%
8
GPT 5.5 Pro Pre Release (xhigh)openai/gpt-5.5-pro-pre-release:xhigh
OpenAI74.23 of 6 benchmarks51.0%39.6%100.0%
9
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI74.1$3.00 / $15.004 of 6 benchmarks72.2%39.0%97.2%1495
10
GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh
OpenAI73.73 of 6 benchmarks51.7%35.4%100.0%
11
GPT-5.5 Pro (xhigh)openai/gpt-5.5-pro:xhigh
OpenAI73.5$30.00 / $180.002 of 6 benchmarks87.7%78.0%
12
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google72.8$1.50 / $9.005 of 6 benchmarks80.0%62.8%26.8%95.6%1507
13
Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh
Qwen72.7$2.00 / $6.003 of 6 benchmarks74.7%46.3%99.4%
14
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google72.6$2.00 / $12.005 of 6 benchmarks88.9%59.6%26.8%95.6%1491
15
GPT-5.4 Pro (xhigh)openai/gpt-5.4-pro:xhigh
OpenAI72.5$30.00 / $180.003 of 6 benchmarks50.0%82.5%58.5%
16
GPT-5.4 (xhigh)openai/gpt-5.4:xhigh
OpenAI72.5$2.50 / $15.004 of 6 benchmarks47.6%78.6%49.0%95.3%
17
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI70.9$5.00 / $30.002 of 6 benchmarks85.3%72.5%
18
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI70.6$1.75 / $14.004 of 6 benchmarks60.0%18.8%96.1%1458
19
Qwen3.7 Maxqwen/qwen3.7-max
Qwen69.9$1.475 / $4.4254 of 6 benchmarks64.6%34.1%95.6%1489
20
Claude Opus 4.6 (64K)anthropic/claude-opus-4.6:64k
Anthropic69.6$5.00 / $25.003 of 6 benchmarks90.0%20.8%94.4%
21
Gemini 3.7 Flash (high)google/gemini-3.7-flash:high
Google69.2$0.375 / $1.8753 of 6 benchmarks71.6%36.6%97.2%
22
GPT-5.2 (xhigh)openai/gpt-5.2:xhigh
OpenAI68.8$1.75 / $14.004 of 6 benchmarks40.7%67.4%31.7%96.1%
23
Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max
Anthropic68.7$5.00 / $25.004 of 6 benchmarks90.0%66.0%26.8%91.1%
24
GPT 5.5 Pro Pre Release (high)openai/gpt-5.5-pro-pre-release:high
OpenAI68.42 of 6 benchmarks52.4%39.6%
25
Grok 4.6 (xhigh)x-ai/grok-4.6:xhigh
SpaceXAI68.1$2.00 / $6.003 of 6 benchmarks66.0%31.7%99.2%
26
Claude Opus 4.6 (32K)anthropic/claude-opus-4.6:32k
Anthropic67.9$5.00 / $25.003 of 6 benchmarks80.0%20.8%93.1%
27
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI67.2$0.5415 / $2.285 of 6 benchmarks39.0%57.2%25.6%96.1%1478
28
DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high
DeepSeek67.0$1.168 / $2.3362 of 6 benchmarks95.6%1469
29
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google66.6$0.75 / $3.754 of 6 benchmarks59.0%21.9%94.2%1513
30
GPT-5.2 (medium)openai/gpt-5.2:medium
OpenAI66.3$1.75 / $14.003 of 6 benchmarks60.0%16.7%93.9%
31
GLM 5.1z-ai/glm-5.1
Z.ai66.2$0.966 / $3.0364 of 6 benchmarks33.5%12.5%93.3%1481
32
Claude Fable 5anthropic/claude-fable-5
Anthropic65.9$10.00 / $50.001 of 6 benchmarks1527
33
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen65.9$0.32 / $1.282 of 6 benchmarks93.3%1471
34
Claude Fable 5 (high)anthropic/claude-fable-5:high
Anthropic65.8$10.00 / $50.001 of 6 benchmarks100.0%
35
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic65.8$5.00 / $25.001 of 6 benchmarks1525
36
GPT-5.2 Pro (xhigh)openai/gpt-5.2-pro:xhigh
OpenAI65.7$21.00 / $168.002 of 6 benchmarks74.0%46.0%
37
Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high
Anthropic65.6$5.00 / $25.001 of 6 benchmarks1516
38
Qwen3.8 Maxqwen/qwen3.8-max
Qwen65.5$2.00 / $6.001 of 6 benchmarks1513
39
GPT-5 (high)openai/gpt-5:high
OpenAI65.5$1.25 / $10.006 of 6 benchmarks32.4%55.4%21.9%91.4%98.1%1435
40
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic65.3$5.00 / $25.003 of 6 benchmarks38.3%14.6%1506
41
Qwen3.6 Plusqwen/qwen3.6-plus
Qwen65.2$0.325 / $1.954 of 6 benchmarks50.0%8.3%93.3%1454
42
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic64.9$5.00 / $25.001 of 6 benchmarks1503
43
GPT-5.5openai/gpt-5.5
OpenAI64.8$5.00 / $30.001 of 6 benchmarks1499
44
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google64.6$0.50 / $3.005 of 6 benchmarks60.0%51.2%17.1%92.8%1476
45
Muse Sparkmeta/muse-spark
Meta64.44 of 6 benchmarks39.0%14.6%88.9%1461
46
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic64.4$5.00 / $25.001 of 6 benchmarks1493
47
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI64.2$5.00 / $30.001 of 6 benchmarks1491
48
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic64.1$5.00 / $25.001 of 6 benchmarks1491
49
Qwen3.6 Max Previewqwen/qwen3.6-max-preview
Qwen64.1$1.027 / $6.1624 of 6 benchmarks50.0%4.2%91.1%1474
50
Claude Fable 5 (low)anthropic/claude-fable-5:low
Anthropic63.8$10.00 / $50.001 of 6 benchmarks97.8%
51
Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low
Anthropic63.8$5.00 / $25.001 of 6 benchmarks97.8%
52
Claude Opus 5anthropic/claude-opus-5
Anthropic63.8$5.00 / $25.001 of 6 benchmarks97.8%
53
Grok 4.6 (high)x-ai/grok-4.6:high
SpaceXAI63.8$2.00 / $6.001 of 6 benchmarks97.8%
54
Muse Spark 1.1meta/muse-spark-1.1
Meta63.6$1.25 / $4.251 of 6 benchmarks1488
55
GPT-5 (medium)openai/gpt-5:medium
OpenAI63.6$1.25 / $10.004 of 6 benchmarks27.2%6.3%87.2%97.9%
56
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai63.6$0.462 / $1.4524 of 6 benchmarks59.2%29.3%86.4%1474
57
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI63.5$0.10 / $0.601 of 6 benchmarks1484
58
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI63.4$1.00 / $6.001 of 6 benchmarks1483
59
Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max
Anthropic63.3$5.00 / $25.003 of 6 benchmarks70.2%31.7%86.7%
60
Hy3tencent/hy3
Tencent63.2$0.132 / $0.5281 of 6 benchmarks1481
61
Inklingthinkingmachines/inkling
Thinking Machines62.9$0.95 / $4.051 of 6 benchmarks1480
62
Grok 4.5x-ai/grok-4.5
SpaceXAI62.8$2.00 / $6.001 of 6 benchmarks1479
63
Gemini 3 Progoogle/gemini-3-pro
Google62.61 of 6 benchmarks1479
64
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI62.5$5.00 / $30.001 of 6 benchmarks1478
65
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Google62.43 of 6 benchmarks37.6%18.8%91.4%
66
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI62.3$1.25 / $10.004 of 6 benchmarks31.0%12.5%88.6%1456
67
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google62.2$1.50 / $9.001 of 6 benchmarks1478
68
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi62.1$0.435 / $0.871 of 6 benchmarks1477
69
Ernie 5.1baidu/ernie-5.1
Baidu61.91 of 6 benchmarks1476
70
Grok 4.5 (high)x-ai/grok-4.5:high
SpaceXAI61.8$2.00 / $6.003 of 6 benchmarks57.2%24.4%97.8%
71
Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high
Google61.6$0.50 / $3.001 of 6 benchmarks95.6%
72
Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high
Google61.6$2.00 / $12.001 of 6 benchmarks95.6%
73
GPT-5.4 (medium)openai/gpt-5.4:medium
OpenAI61.6$2.50 / $15.001 of 6 benchmarks95.6%
74
GPT-5.6 Sol (low)openai/gpt-5.6-sol:low
OpenAI61.6$5.00 / $30.001 of 6 benchmarks95.6%
75
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen61.4$0.39 / $2.342 of 6 benchmarks88.9%1449
76
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic61.4$5.00 / $25.001 of 6 benchmarks1472
77
Gemma 4 31Bgoogle/gemma-4-31b-it
Google61.2$0.10 / $0.341 of 6 benchmarks1472
78
Qwen3.5 Max Previewqwen/qwen3.5-max-preview
Qwen60.91 of 6 benchmarks1470
79
Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06
Google60.9$1.25 / $10.001 of 6 benchmarks95.9%
80
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI60.8$0.57 / $2.851 of 6 benchmarks1470
81
Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh
Anthropic60.7$2.00 / $10.001 of 6 benchmarks94.7%
82
Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k
Anthropic60.7$5.00 / $25.001 of 6 benchmarks1469
83
DeepSeek V4 Flash 0731 (max)deepseek/deepseek-v4-flash-0731:max
DeepSeek60.4$0.14 / $0.283 of 6 benchmarks57.5%24.4%94.4%
84
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic60.4$2.00 / $10.001 of 6 benchmarks1468
85
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google60.2$0.12 / $0.401 of 6 benchmarks1468
86
Qwen3 Maxqwen/qwen3-max
Qwen60.1$0.78 / $3.903 of 6 benchmarks73.3%97.1%1426
87
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI60.11 of 6 benchmarks1467
88
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic59.8$3.00 / $15.001 of 6 benchmarks1463
89
GPT-5 Pro (high)openai/gpt-5-pro:high
OpenAI59.6$15.00 / $120.003 of 6 benchmarks60.0%55.8%19.5%
90
Claude Opus 5 (low)anthropic/claude-opus-5:low
Anthropic59.5$5.00 / $25.001 of 6 benchmarks93.3%
91
Kimi K3 (high)moonshotai/kimi-k3:high
MoonshotAI59.5$3.00 / $15.001 of 6 benchmarks93.3%
92
GPT-5.4openai/gpt-5.4
OpenAI59.4$2.50 / $15.001 of 6 benchmarks1460
93
Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k
Anthropic58.9$3.00 / $15.001 of 6 benchmarks1455
94
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google58.6$0.50 / $3.001 of 6 benchmarks1453
95
Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max
Anthropic58.6$2.00 / $10.003 of 6 benchmarks65.6%29.3%80.0%
96
Muse Glimmer 30Bmeta/muse-glimmer-30b
Meta58.4$0.35 / $1.501 of 6 benchmarks1453
97
Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309
xAI58.41 of 6 benchmarks1453
98
Mimo v2 Proxiaomi/mimo-v2-pro
Xiaomi58.21 of 6 benchmarks1452
99
GPT-5.2 Chatopenai/gpt-5.2-chat
OpenAI58.1$1.75 / $14.001 of 6 benchmarks1452
100
Dola Seed 2.0 Probytedance/dola-seed-2.0-pro
ByteDance57.91 of 6 benchmarks1451
101
GPT-5 Mini (high)openai/gpt-5-mini:high
OpenAI57.9$0.25 / $2.006 of 6 benchmarks27.2%46.7%12.2%86.7%97.8%1405
102
Qwen3.6 27Bqwen/qwen3.6-27b
Qwen57.9$0.60 / $3.601 of 6 benchmarks91.1%
103
o3openai/o3
OpenAI57.6$2.00 / $8.001 of 6 benchmarks1447
104
Kimi K2p5moonshotai/kimi-k2p5
Moonshot AI57.63 of 6 benchmarks27.9%4.2%92.2%
105
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI57.6$0.75 / $4.504 of 6 benchmarks50.0%2.1%87.2%1439
106
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek57.5$1.168 / $2.3361 of 6 benchmarks1445
107
Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k
Anthropic57.4$15.00 / $75.001 of 6 benchmarks1444
108
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai57.2$1.20 / $4.001 of 6 benchmarks1443
109
Grok 4.1 (thinking)x-ai/grok-4.1:thinking
xAI56.91 of 6 benchmarks1442
110
Gemini 3.5 Flash (low)google/gemini-3.5-flash:low
Google56.8$1.50 / $9.001 of 6 benchmarks88.9%
111
GPT-5.6 Terra (low)openai/gpt-5.6-terra:low
OpenAI56.8$1.00 / $6.001 of 6 benchmarks88.9%
112
gpt-oss-120b (high)openai/gpt-oss-120b:high
OpenAI56.8$0.03 / $0.171 of 6 benchmarks88.9%
113
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
NVIDIA56.8$0.60 / $3.601 of 6 benchmarks1442
114
GPT-5 Mini (medium)openai/gpt-5-mini:medium
OpenAI56.8$0.25 / $2.004 of 6 benchmarks20.3%4.2%78.3%96.8%
115
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi56.6$0.14 / $0.281 of 6 benchmarks1441
116
Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo
Moonshot AI56.53 of 6 benchmarks20.0%83.1%1436
117
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek56.4$0.0643 / $0.12851 of 6 benchmarks1441
118
Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh
Anthropic56.3$5.00 / $25.003 of 6 benchmarks43.8%0.0%97.8%
119
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI56.11 of 6 benchmarks1440
120
Ernie 5.0 0110baidu/ernie-5.0-0110
Baidu55.81 of 6 benchmarks1437
121
o1 (medium)openai/o1:medium
OpenAI55.7$15.00 / $60.002 of 6 benchmarks73.3%94.4%
122
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google55.6$0.25 / $1.501 of 6 benchmarks1437
123
GPT-5.2 (low)openai/gpt-5.2:low
OpenAI55.5$1.75 / $14.003 of 6 benchmarks40.0%6.3%78.9%
124
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI55.3$0.71 / $3.503 of 6 benchmarks54.0%12.2%95.6%
125
Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b
Qwen55.3$0.15 / $1.001 of 6 benchmarks86.7%
126
Qwen3.7 Flashqwen/qwen3.7-flash
Qwen55.3$0.03 / $0.131 of 6 benchmarks86.7%
127
Mimo v2 Omnixiaomi/mimo-v2-omni
Xiaomi55.21 of 6 benchmarks1435
128
GPT-5.2openai/gpt-5.2
OpenAI54.9$1.75 / $14.001 of 6 benchmarks1432
129
o3 (high)openai/o3:high
OpenAI54.9$2.00 / $8.004 of 6 benchmarks18.7%2.1%83.9%97.8%
130
Qwen3.5 Plusqwen/qwen3.5-plus
Qwen54.8$0.30 / $1.803 of 6 benchmarks50.0%2.1%86.7%
131
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral54.8$1.50 / $7.501 of 6 benchmarks1429
132
DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking
DeepSeek54.6$0.27 / $0.411 of 6 benchmarks1429
133
Claude Sonnet 4.6 (32K)anthropic/claude-sonnet-4.6:32k
Anthropic54.4$3.00 / $15.001 of 6 benchmarks85.8%
134
o4 Mini Highopenai/o4-mini-high
OpenAI54.4$1.10 / $4.405 of 6 benchmarks24.8%36.1%4.9%81.7%97.8%
135
Grok 4.1x-ai/grok-4.1
xAI54.41 of 6 benchmarks1428
136
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen54.2$0.195 / $1.561 of 6 benchmarks1428
137
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI53.7$1.25 / $10.003 of 6 benchmarks26.9%4.2%85.6%
138
MiniMax M3minimax/minimax-m3
MiniMax53.6$0.30 / $1.202 of 6 benchmarks71.1%1440
139
GPT-5.3 Chatopenai/gpt-5.3-chat
OpenAI53.61 of 6 benchmarks1427
140
Claude Opus 4.8 (none)anthropic/claude-opus-4.8:none
Anthropic53.5$5.00 / $25.001 of 6 benchmarks84.4%
141
GPT-5.4 (low)openai/gpt-5.4:low
OpenAI53.5$2.50 / $15.001 of 6 benchmarks84.4%
142
GPT-5.5 (low)openai/gpt-5.5:low
OpenAI53.5$5.00 / $30.001 of 6 benchmarks84.4%
143
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash
DeepSeek53.3$0.0643 / $0.12851 of 6 benchmarks1425
144
Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview
Tencent53.21 of 6 benchmarks1425
145
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek53.1$0.269 / $0.401 of 6 benchmarks1424
146
Inkling Small (xhigh)thinkingmachines/inkling-small:xhigh
Thinking Machines53.0$0.45 / $1.203 of 6 benchmarks46.3%17.1%90.0%
147
GPT-5.4 Nano (high)openai/gpt-5.4-nano:high
OpenAI52.9$0.20 / $1.255 of 6 benchmarks25.9%44.9%12.2%87.8%1423
148
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google52.9$0.30 / $2.501 of 6 benchmarks1423
149
o3 (medium)openai/o3:medium
OpenAI52.6$2.00 / $8.002 of 6 benchmarks16.9%84.4%
150
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen52.6$0.29 / $2.401 of 6 benchmarks1423
151
Grok 4.20 0309 (reasoning)x-ai/grok-4.20-0309:reasoning
xAI52.63 of 6 benchmarks44.9%17.1%92.2%
152
GPT-5.1openai/gpt-5.1
OpenAI52.5$1.25 / $10.001 of 6 benchmarks1423
153
o1 (high)openai/o1:high
OpenAI52.4$15.00 / $60.003 of 6 benchmarks9.3%73.3%94.7%
154
MiniMax M2.7minimax/minimax-m2.7
MiniMax52.3$0.30 / $1.201 of 6 benchmarks1423
155
Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium
Anthropic52.3$3.00 / $15.001 of 6 benchmarks82.2%
156
Gemini 3.6 Flash (low)google/gemini-3.6-flash:low
Google52.3$0.75 / $3.751 of 6 benchmarks82.2%
157
Qwen3.5 397B A17B (none)qwen/qwen3.5-397b-a17b:none
Qwen52.3$0.39 / $2.341 of 6 benchmarks82.2%
158
Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k
Anthropic52.2$15.00 / $75.001 of 6 benchmarks1421
159
o3 Mini (medium)openai/o3-mini:medium
OpenAI52.1$1.10 / $4.403 of 6 benchmarks11.3%63.9%95.2%
160
Grok 4.3x-ai/grok-4.3
SpaceXAI51.9$1.25 / $2.501 of 6 benchmarks1419
161
Grok 4.3 (high)x-ai/grok-4.3:high
SpaceXAI51.8$1.25 / $2.503 of 6 benchmarks42.8%14.6%93.3%
162
Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507
Qwen51.8$0.09 / $0.551 of 6 benchmarks1418
163
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI51.61 of 6 benchmarks1418
164
Kimi K2 0905moonshotai/kimi-k2-0905
MoonshotAI51.5$0.60 / $2.501 of 6 benchmarks1417
165
GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh
OpenAI51.4$0.75 / $4.503 of 6 benchmarks51.2%9.8%88.9%
166
Claude Opus 4.5 (16K)anthropic/claude-opus-4.5:16k
Anthropic51.4$5.00 / $25.003 of 6 benchmarks40.0%2.1%81.7%
167
Claude Opus 4.5anthropic/claude-opus-4.5
Anthropic51.4$5.00 / $25.004 of 6 benchmarks20.7%4.2%48.1%1465
168
Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct
Qwen51.3$0.10 / $1.101 of 6 benchmarks1417
169
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp
DeepSeek51.2$0.27 / $0.411 of 6 benchmarks1416
170
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek51.2$0.269 / $0.403 of 6 benchmarks22.1%2.1%1429
171
Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high
Google51.1$0.25 / $1.501 of 6 benchmarks80.0%
172
Gemini 3.5 Flash (minimal)google/gemini-3.5-flash:minimal
Google51.1$1.50 / $9.001 of 6 benchmarks80.0%
173
Gemini 3.6 Flash (minimal)google/gemini-3.6-flash:minimal
Google51.1$0.75 / $3.751 of 6 benchmarks80.0%
174
Qwen3.7 Plus (none)qwen/qwen3.7-plus:none
Qwen51.1$0.32 / $1.281 of 6 benchmarks80.0%
175
R1deepseek/deepseek-r1
DeepSeek51.1$0.70 / $2.503 of 6 benchmarks53.3%93.0%1412
176
o4 Miniopenai/o4-mini
OpenAI51.1$1.10 / $4.401 of 6 benchmarks1415
177
GLM 5z-ai/glm-5
Z.ai50.9$0.60 / $1.924 of 6 benchmarks16.4%2.1%80.0%1443
178
LongCat Flash Chatmeituan/longcat-flash-chat
Meituan50.91 of 6 benchmarks1415
179
DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking
DeepSeek50.8$0.25 / $0.951 of 6 benchmarks1414
180
DeepSeek V3.1deepseek/deepseek-chat-v3.1
DeepSeek50.6$0.25 / $0.951 of 6 benchmarks1413
181
Claude Sonnet 4.6 (16K)anthropic/claude-sonnet-4.6:16k
Anthropic50.6$3.00 / $15.002 of 6 benchmarks80.0%0.0%
182
GPT 5 Chatopenai/gpt-5-chat
OpenAI50.31 of 6 benchmarks1413
183
DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max
DeepSeek50.2$1.168 / $2.3363 of 6 benchmarks45.3%2.4%96.7%
184
Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning
xAI50.11 of 6 benchmarks1410
185
Qwen3.7 Flash (none)qwen/qwen3.7-flash:none
Qwen50.0$0.03 / $0.131 of 6 benchmarks77.8%
186
Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203
Baidu49.91 of 6 benchmarks1409
187
Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct
Qwen49.6$0.26 / $1.041 of 6 benchmarks1408
188
Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high
Anthropic49.5$3.00 / $15.001 of 6 benchmarks75.6%
189
Step 3.5 Flashstepfun/step-3.5-flash
StepFun49.4$0.10 / $0.301 of 6 benchmarks1406
190
GPT-5 Nano (medium)openai/gpt-5-nano:medium
OpenAI49.3$0.05 / $0.404 of 6 benchmarks10.0%2.1%74.2%95.2%
191
Claude Sonnet 4.5 (59K)anthropic/claude-sonnet-4.5:59k
Anthropic49.1$3.00 / $15.002 of 6 benchmarks13.5%77.8%
192
Gemma 4 31B (minimal)google/gemma-4-31b-it:minimal
Google48.9$0.10 / $0.341 of 6 benchmarks73.3%
193
Mistral Large 3mistralai/mistral-large-3
Mistral AI48.81 of 6 benchmarks1404
194
Chatgpt 4oopenai/chatgpt-4o
OpenAI48.61 of 6 benchmarks1404
195
Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking
Qwen48.5$0.40 / $4.001 of 6 benchmarks1404
196
Claude Sonnet 4 (32K)anthropic/claude-sonnet-4:32k
Anthropic48.1$3.00 / $15.001 of 6 benchmarks71.1%
197
Claude Sonnet 4.5 (16K)anthropic/claude-sonnet-4.5:16k
Anthropic48.1$3.00 / $15.001 of 6 benchmarks71.1%
198
Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max
Anthropic48.1$3.00 / $15.001 of 6 benchmarks71.1%
199
Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k
Anthropic48.1$3.00 / $15.001 of 6 benchmarks1403
200
Grok 4 0709x-ai/grok-4-0709
xAI48.03 of 6 benchmarks19.7%2.1%1427
201
Hunyuan T1tencent/hunyuan-t1
Tencent47.91 of 6 benchmarks1401
202
Claude Opus 4.5 (32K)anthropic/claude-opus-4.5:32k
Anthropic47.9$5.00 / $25.004 of 6 benchmarks20.7%34.4%4.9%86.1%
203
Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview
Google47.9$1.25 / $10.002 of 6 benchmarks30.0%2.1%
204
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Qwen47.8$0.225 / $1.801 of 6 benchmarks1400
205
Gemini 2.5 Progoogle/gemini-2.5-pro
Google47.7$1.25 / $10.006 of 6 benchmarks14.1%24.6%0.0%84.7%95.6%1441
206
Qwen3 32Bqwen/qwen3-32b
Qwen47.6$0.08 / $0.281 of 6 benchmarks1399
207
Claude Sonnet 4.5 (32K)anthropic/claude-sonnet-4.5:32k
Anthropic47.5$3.00 / $15.005 of 6 benchmarks15.2%23.9%2.4%77.8%97.7%
208
Claude Haiku 4.5 (32K)anthropic/claude-haiku-4.5:32k
Anthropic47.4$1.00 / $5.004 of 6 benchmarks5.9%2.1%66.7%96.4%
209
Mistral Medium 2508mistralai/mistral-medium-2508
Mistral AI47.31 of 6 benchmarks1398
210
Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022
Baidu47.11 of 6 benchmarks1396
211
Kimi K3 (low)moonshotai/kimi-k3:low
MoonshotAI47.0$3.00 / $15.001 of 6 benchmarks68.9%
212
GPT-5.4 Nano (low)openai/gpt-5.4-nano:low
OpenAI47.0$0.20 / $1.251 of 6 benchmarks68.9%
213
GPT-5.6 Sol (none)openai/gpt-5.6-sol:none
OpenAI47.0$5.00 / $30.001 of 6 benchmarks68.9%
214
Qwen3.6 35B A3B (none)qwen/qwen3.6-35b-a3b:none
Qwen47.0$0.15 / $1.001 of 6 benchmarks68.9%
215
MiniMax M2.5minimax/minimax-m2.5
MiniMax46.9$0.22 / $0.901 of 6 benchmarks1396
216
GPT-5.1 (low)openai/gpt-5.1:low
OpenAI46.7$1.25 / $10.002 of 6 benchmarks17.3%63.9%
217
DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus
DeepSeek46.6$0.27 / $0.951 of 6 benchmarks1394
218
Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking
Qwen46.5$0.455 / $1.821 of 6 benchmarks1393
219
Inkling (xhigh)thinkingmachines/inkling:xhigh
Thinking Machines46.4$0.95 / $4.053 of 6 benchmarks33.3%4.9%88.9%
220
R1 0528deepseek/deepseek-r1-0528
DeepSeek46.3$0.50 / $2.154 of 6 benchmarks0.0%66.4%96.6%1395
221
Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking
Qwen46.2$0.15 / $1.201 of 6 benchmarks1391
222
GPT-5.6 Luna (low)openai/gpt-5.6-luna:low
OpenAI46.2$0.10 / $0.601 of 6 benchmarks66.7%
223
Qwen3.6 27B (none)qwen/qwen3.6-27b:none
Qwen46.2$0.60 / $3.601 of 6 benchmarks66.7%
224
MiniMax M2.1minimax/minimax-m2.1
MiniMax46.1$0.30 / $1.201 of 6 benchmarks1390
225
GLM 4.5 Airz-ai/glm-4.5-air
Z.ai45.9$0.13 / $0.851 of 6 benchmarks1390
226
Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507
Qwen45.8$0.23 / $2.304 of 6 benchmarks20.0%0.0%86.7%1397
227
Grok 3 Mini Beta (low)x-ai/grok-3-mini-beta:low
xAI45.73 of 6 benchmarks2.8%62.2%90.9%
228
Kimi K2 0711moonshotai/kimi-k2
MoonshotAI45.6$0.57 / $2.301 of 6 benchmarks1388
229
Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA45.31 of 6 benchmarks1386
230
Qwen3.6 Flashqwen/qwen3.6-flash
Qwen45.2$0.1875 / $1.1253 of 6 benchmarks20.0%0.0%84.4%
231
Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k
Anthropic45.21 of 6 benchmarks1385
232
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Arcee AI45.1$0.22 / $0.851 of 6 benchmarks1384
233
GPT-5.2 (none)openai/gpt-5.2:none
OpenAI45.0$1.75 / $14.001 of 6 benchmarks62.2%
234
INTELLECT-3prime-intellect/intellect-3
Prime Intellect44.91 of 6 benchmarks1382
235
o3 Miniopenai/o3-mini
OpenAI44.8$1.10 / $4.401 of 6 benchmarks1382
236
o4 Mini (medium)openai/o4-mini:medium
OpenAI44.7$1.10 / $4.403 of 6 benchmarks19.0%2.1%73.3%
237
Claude 3.7 Sonnet (32K)anthropic/claude-3-7-sonnet:32k
Anthropic44.73 of 6 benchmarks3.5%53.3%90.0%
238
gpt-oss-120bopenai/gpt-oss-120b
OpenAI44.6$0.03 / $0.171 of 6 benchmarks1381
239
GPT 5.5 Instantopenai/gpt-5.5-instant
OpenAI44.64 of 6 benchmarks26.3%2.4%68.1%1462
240
Claude Opus 4 (16K)anthropic/claude-opus-4:16k
Anthropic44.6$15.00 / $75.001 of 6 benchmarks60.0%
241
Gemini 3.5 Flash Lite (low)google/gemini-3.5-flash-lite:low
Google44.6$0.30 / $2.501 of 6 benchmarks60.0%
242
Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1
NVIDIA44.51 of 6 benchmarks1380
243
Claude Opus 4.1anthropic/claude-opus-4.1
Anthropic44.3$15.00 / $75.003 of 6 benchmarks5.9%40.0%1434
244
Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507
Qwen44.3$0.0482 / $0.19311 of 6 benchmarks1380
245
Mimo v2 Flashxiaomi/mimo-v2-flash
Xiaomi44.21 of 6 benchmarks1377
246
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google44.1$0.30 / $2.504 of 6 benchmarks4.8%4.2%70.8%1406
247
Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b
NVIDIA44.0$0.085 / $0.401 of 6 benchmarks1376
248
GPT-5.4 (none)openai/gpt-5.4:none
OpenAI44.0$2.50 / $15.001 of 6 benchmarks57.8%
249
GPT-5.5 (none)openai/gpt-5.5:none
OpenAI44.0$5.00 / $30.001 of 6 benchmarks57.8%
250
o1openai/o1
OpenAI43.9$15.00 / $60.003 of 6 benchmarks31.1%81.7%1409
251
GPT-5 Nano (high)openai/gpt-5-nano:high
OpenAI43.9$0.05 / $0.406 of 6 benchmarks20.0%20.0%2.4%81.1%94.9%1346
252
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen43.91 of 6 benchmarks1376
253
o4 Mini (low)openai/o4-mini:low
OpenAI43.9$1.10 / $4.402 of 6 benchmarks10.7%57.8%
254
GPT 4.5 Previewopenai/gpt-4.5-preview
OpenAI43.93 of 6 benchmarks37.8%78.6%1408
255
Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking
Xiaomi43.81 of 6 benchmarks1374
256
Claude Opus 4.1 (27K)anthropic/claude-opus-4.1:27k
Anthropic43.7$15.00 / $75.003 of 6 benchmarks7.2%4.2%68.9%
257
GPT-5 Mini (minimal)openai/gpt-5-mini:minimal
OpenAI43.6$0.25 / $2.001 of 6 benchmarks55.6%
258
o3 (low)openai/o3:low
OpenAI43.4$2.00 / $8.002 of 6 benchmarks9.7%60.0%
259
MiniMax M1minimax/minimax-m1
MiniMax43.3$0.55 / $2.201 of 6 benchmarks1371
260
gpt-oss-20b (high)openai/gpt-oss-20b:high
OpenAI43.3$0.03 / $0.131 of 6 benchmarks53.9%
261
Grok 3 Mini Betax-ai/grok-3-mini-beta
xAI43.01 of 6 benchmarks1368
262
Claude 3.7 Sonnet (16K)anthropic/claude-3-7-sonnet:16k
Anthropic43.03 of 6 benchmarks4.1%46.7%86.3%
263
Qwen3.5-Flashqwen/qwen3.5-flash-02-23
Qwen43.0$0.065 / $0.264 of 6 benchmarks10.0%0.0%84.4%1403
264
GLM 4.7 Flashz-ai/glm-4.7-flash
Z.ai42.9$0.06 / $0.401 of 6 benchmarks1366
265
GLM 4.6z-ai/glm-4.6
Z.ai42.9$0.55 / $2.203 of 6 benchmarks3.8%2.1%1420
266
Claude Sonnet 4 (16K)anthropic/claude-sonnet-4:16k
Anthropic42.9$3.00 / $15.001 of 6 benchmarks53.3%
267
GPT-5.6 Terra (none)openai/gpt-5.6-terra:none
OpenAI42.9$1.00 / $6.001 of 6 benchmarks53.3%
268
o1 (low)openai/o1:low
OpenAI42.9$15.00 / $60.001 of 6 benchmarks53.3%
269
Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking
Google42.8$0.10 / $0.401 of 6 benchmarks1365
270
o3 Mini Highopenai/o3-mini-high
OpenAI42.8$1.10 / $4.406 of 6 benchmarks12.4%18.6%0.0%76.9%96.5%1405
271
Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high
xAI42.65 of 6 benchmarks5.9%0.0%77.8%88.1%1387
272
QwQ 32Bqwen/qwq-32b
Qwen42.61 of 6 benchmarks1364
273
Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking
Google42.5$0.10 / $0.401 of 6 benchmarks1364
274
Gemini 3.5 Flash Lite (minimal)google/gemini-3.5-flash-lite:minimal
Google42.4$0.30 / $2.501 of 6 benchmarks51.1%
275
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI42.4$0.40 / $1.604 of 6 benchmarks10.0%44.7%87.3%1354
276
Qwen 2.5 (max)qwen/qwen-2.5:max
Qwen42.31 of 6 benchmarks1363
277
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic42.3$1.00 / $5.004 of 6 benchmarks4.1%35.8%86.9%1398
278
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
MoonshotAI42.3$0.60 / $2.502 of 6 benchmarks21.4%0.0%
279
O1 Mini (high)openai/o1-mini:high
OpenAI42.23 of 6 benchmarks1.4%46.9%89.2%
280
GLM 4.7z-ai/glm-4.7
Z.ai42.1$0.40 / $1.754 of 6 benchmarks2.4%0.0%83.3%1428
281
Step 3stepfun/step-3
StepFun42.01 of 6 benchmarks1362
282
O1 Miniopenai/o1-mini
OpenAI41.91 of 6 benchmarks1362
283
Trinity Large Previewarcee-ai/trinity-large-preview
Arcee AI41.81 of 6 benchmarks1362
284
GLM 4.5Vz-ai/glm-4.5v
Z.ai41.6$0.60 / $1.801 of 6 benchmarks1360
285
DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none
DeepSeek41.4$1.168 / $2.3361 of 6 benchmarks46.7%
286
GPT-5 Nano (low)openai/gpt-5-nano:low
OpenAI41.4$0.05 / $0.401 of 6 benchmarks46.7%
287
GPT-5 (minimal)openai/gpt-5:minimal
OpenAI41.4$1.25 / $10.001 of 6 benchmarks46.7%
288
GPT-5.4 Nano (none)openai/gpt-5.4-nano:none
OpenAI41.4$0.20 / $1.251 of 6 benchmarks46.7%
289
MiniMax M2minimax/minimax-m2
MiniMax41.3$0.255 / $1.021 of 6 benchmarks1354
290
Claude Opus 4 (27K)anthropic/claude-opus-4:27k
Anthropic41.3$15.00 / $75.003 of 6 benchmarks4.1%4.2%64.4%
291
Qwen3 30B A3Bqwen/qwen3-30b-a3b
Qwen41.0$0.12 / $0.501 of 6 benchmarks1352
292
Ling Flash 2.0inclusionai/ling-flash-2.0
inclusionAI40.91 of 6 benchmarks1352
293
Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b
NVIDIA40.6$0.05 / $0.201 of 6 benchmarks1351
294
Gemini 3.1 Flash Lite (low)google/gemini-3.1-flash-lite:low
Google40.6$0.25 / $1.501 of 6 benchmarks44.4%
295
o3 Mini (low)openai/o3-mini:low
OpenAI40.6$1.10 / $4.401 of 6 benchmarks44.4%
296
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic40.3$3.00 / $15.004 of 6 benchmarks9.3%2.1%35.6%1427
297
Hunyuan TurboStencent/hunyuan-turbos
Tencent40.31 of 6 benchmarks1347
298
GPT-5.6 Luna (none)openai/gpt-5.6-luna:none
OpenAI40.1$0.10 / $0.601 of 6 benchmarks40.0%
299
Ring Flash 2.0inclusionai/ring-flash-2.0
inclusionAI40.11 of 6 benchmarks1340
300
Mistral Small 2506mistralai/mistral-small-2506
Mistral AI39.81 of 6 benchmarks1338
301
O1 Mini (medium)openai/o1-mini:medium
OpenAI39.73 of 6 benchmarks1.7%44.7%84.3%
302
Claude 3.7 Sonnet (64K)anthropic/claude-3-7-sonnet:64k
Anthropic39.64 of 6 benchmarks3.1%0.0%57.8%91.2%
303
gpt-oss-20bopenai/gpt-oss-20b
OpenAI39.6$0.03 / $0.131 of 6 benchmarks1336
304
Nova 2 Liteamazon/nova-2-lite-v1
Amazon39.5$0.30 / $2.501 of 6 benchmarks1333
305
Gemini 3.1 Flash Lite (minimal)google/gemini-3.1-flash-lite:minimal
Google39.5$0.25 / $1.501 of 6 benchmarks37.8%
306
Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05
Google39.31 of 6 benchmarks1326
307
Granite 4.1 8Bibm-granite/granite-4.1-8b
IBM39.0$0.05 / $0.101 of 6 benchmarks1318
308
Gemma 3 12Bgoogle/gemma-3-12b-it
Google38.9$0.05 / $0.151 of 6 benchmarks1318
309
GPT-5 Nano (minimal)openai/gpt-5-nano:minimal
OpenAI38.8$0.05 / $0.401 of 6 benchmarks35.6%
310
OLMo 3 32B Thinkallenai/olmo-3-32b-think
Allen Institute for AI38.51 of 6 benchmarks1311
311
Claude Opus 4anthropic/claude-opus-4
Anthropic38.4$15.00 / $75.005 of 6 benchmarks4.5%0.0%42.2%85.0%1404
312
Command A (03-2025)cohere/command-a-03-2025
Cohere38.21 of 6 benchmarks1309
313
Grok 3 Betax-ai/grok-3-beta
xAI37.95 of 6 benchmarks3.8%0.0%55.6%88.8%1373
314
GLM 5.2 (none)z-ai/glm-5.2:none
Z.ai37.8$0.462 / $1.4521 of 6 benchmarks28.9%
315
Claude Sonnet 4 (59K)anthropic/claude-sonnet-4:59k
Anthropic37.6$3.00 / $15.002 of 6 benchmarks0.0%68.9%
316
OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct
Allen Institute for AI37.61 of 6 benchmarks1304
317
GPT-5.4 Mini (none)openai/gpt-5.4-mini:none
OpenAI37.5$0.75 / $4.501 of 6 benchmarks26.7%
318
Step 1o Turbo 202506stepfun/step-1o-turbo-202506
StepFun37.21 of 6 benchmarks1297
319
Qwen3 235B A22Bqwen/qwen3-235b-a22b
Qwen36.8$0.455 / $1.823 of 6 benchmarks0.0%68.9%1392
320
OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think
Allen Institute for AI36.81 of 6 benchmarks1296
321
Magistral Medium 2506mistralai/magistral-medium-2506
Mistral AI36.01 of 6 benchmarks1285
322
GPT-4.1openai/gpt-4.1
OpenAI35.9$2.00 / $8.005 of 6 benchmarks5.5%0.0%38.3%83.0%1374
323
Gemma 3 27Bgoogle/gemma-3-27b-it
Google35.8$0.08 / $0.453 of 6 benchmarks22.2%74.0%1323
324
Hunyuan Large Visiontencent/hunyuan-large-vision
Tencent35.81 of 6 benchmarks1280
325
Gemini 1.5 Pro 002google/gemini-1.5-pro-002
Google35.73 of 6 benchmarks23.1%70.4%1339
326
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic35.7$3.00 / $15.005 of 6 benchmarks4.1%0.0%28.9%84.4%1389
327
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google35.74 of 6 benchmarks1.7%31.1%82.2%1356
328
Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct
Qwen35.21 of 6 benchmarks1270
329
Claude Opus 4.1 (16K)anthropic/claude-opus-4.1:16k
Anthropic34.9$15.00 / $75.002 of 6 benchmarks0.0%64.4%
330
Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0
Amazon34.91 of 6 benchmarks1269
331
Gemma 3n E4Bgoogle/gemma-3n-e4b-it
Google34.61 of 6 benchmarks1260
332
GPT-5.1 (none)openai/gpt-5.1:none
OpenAI34.5$1.25 / $10.002 of 6 benchmarks2.1%37.8%
333
Gemma 3 4Bgoogle/gemma-3-4b-it
Google34.31 of 6 benchmarks1254
334
Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0
Amazon34.01 of 6 benchmarks1244
335
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic33.94 of 6 benchmarks3.1%21.9%68.2%1363
336
Command R+ (08-2024)cohere/command-r-plus-08-2024
Cohere33.8$2.50 / $10.001 of 6 benchmarks1231
337
DeepSeek V3 0324deepseek/deepseek-chat-v3-0324
DeepSeek33.6$0.27 / $1.124 of 6 benchmarks0.0%37.8%75.5%1369
338
Mistral Medium 2505mistralai/mistral-medium-2505
Mistral AI33.64 of 6 benchmarks0.3%32.2%81.6%1348
339
OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct
Allen Institute for AI33.31 of 6 benchmarks1227
340
DeepSeek V3deepseek/deepseek-chat
DeepSeek33.2$0.2574 / $1.02874 of 6 benchmarks1.7%48.9%64.8%1311
341
Claude Opus 4.1 (32K)anthropic/claude-opus-4.1:32k
Anthropic33.2$15.00 / $75.002 of 6 benchmarks12.6%2.4%
342
Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0
Amazon33.21 of 6 benchmarks1224
343
QwQ 32B Previewqwen/qwq-32b-preview
Qwen33.01 of 6 benchmarks1210
344
Command R (08-2024)cohere/command-r-08-2024
Cohere32.9$0.15 / $0.601 of 6 benchmarks1207
345
Gemini 3.5 Flash Lite (high)google/gemini-3.5-flash-lite:high
Google32.9$0.30 / $2.503 of 6 benchmarks26.0%0.0%71.1%
346
Mistral Nemomistralai/mistral-nemo
Mistral32.8$0.019 / $0.031 of 6 benchmarks10.8%
347
GLM 4.5z-ai/glm-4.5
Z.ai32.7$0.60 / $2.203 of 6 benchmarks0.0%0.0%1413
348
Qwen (max)qwen/qwen:max
Qwen32.63 of 6 benchmarks1.0%16.1%67.2%
349
Grok 2 1212x-ai/grok-2-1212
xAI31.53 of 6 benchmarks0.7%11.5%63.5%
350
GPT 4 1106 Previewopenai/gpt-4-1106-preview
OpenAI31.32 of 6 benchmarks40.0%1303
351
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta31.14 of 6 benchmarks0.7%20.6%73.0%1317
352
Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct
Qwen31.1$0.36 / $0.403 of 6 benchmarks8.1%63.2%1296
353
GPT-4.1 Nanoopenai/gpt-4.1-nano
OpenAI29.8$0.10 / $0.404 of 6 benchmarks1.0%28.9%70.0%1274
354
Magistral Small 2506mistralai/magistral-small-2506
Mistral AI29.32 of 6 benchmarks0.0%30.0%
355
GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13
OpenAI29.2$5.00 / $15.003 of 6 benchmarks6.3%51.0%1305
356
GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18
OpenAI28.5$0.15 / $0.603 of 6 benchmarks6.9%52.6%1276
357
GPT-4 Turboopenai/gpt-4-turbo
OpenAI27.9$10.00 / $30.003 of 6 benchmarks6.7%46.7%1296
358
Mistral Large 2407mistralai/mistral-large-2407
Mistral27.4$2.00 / $6.003 of 6 benchmarks8.5%44.8%1288
359
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20
OpenAI27.2$2.50 / $10.003 of 6 benchmarks0.3%6.3%49.8%
360
Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503
Mistral AI26.93 of 6 benchmarks5.8%46.8%1278
361
Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct
Mistral26.9$2.00 / $6.002 of 6 benchmarks24.2%1228
362
Gemini 1.5 Pro 001google/gemini-1.5-pro-001
Google26.73 of 6 benchmarks6.8%40.8%1299
363
Claude 3.5 Sonnetanthropic/claude-3-5-sonnet
Anthropic26.75 of 6 benchmarks1.0%0.0%8.5%57.0%1351
364
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta26.54 of 6 benchmarks0.0%7.8%62.3%1308
365
GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06
OpenAI26.4$2.50 / $10.004 of 6 benchmarks0.3%6.4%53.3%1309
366
Claude 3 Opusanthropic/claude-3-opus
Anthropic26.23 of 6 benchmarks4.7%37.5%1312
367
Gemini 1.5 Flash 002google/gemini-1.5-flash-002
Google26.14 of 6 benchmarks0.0%16.3%61.9%1289
368
Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001
Google26.02 of 6 benchmarks4.6%1230
369
Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct
Meta25.9$0.10 / $0.323 of 6 benchmarks5.1%41.6%1296
370
Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501
Mistral AI25.33 of 6 benchmarks5.3%44.8%1262
371
Claude 2anthropic/claude-2
Anthropic25.12 of 6 benchmarks2.5%11.7%
372
Mistral Large 2411mistralai/mistral-large-2411
Mistral AI25.04 of 6 benchmarks0.3%7.8%50.3%1282
373
Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct
Meta23.4$0.40 / $0.403 of 6 benchmarks3.6%36.7%1269
374
Claude 3.5 Haikuanthropic/claude-3-5-haiku
Anthropic23.24 of 6 benchmarks0.3%4.3%46.4%1286
375
Gemini 1.5 Flash 001google/gemini-1.5-flash-001
Google22.83 of 6 benchmarks3.9%25.1%1258
376
Claude 3 Sonnetanthropic/claude-3-sonnet
Anthropic21.53 of 6 benchmarks2.5%18.2%1253
377
Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct
Meta21.0$0.05 / $0.083 of 6 benchmarks2.5%22.9%1189
378
Claude 3 Haikuanthropic/claude-3-haiku
Anthropic20.9$0.25 / $1.253 of 6 benchmarks1.8%14.9%1231
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • Epoch AI Benchmarking HubCC BY 4.0

    Benchmark runs by Epoch AI, from the AI Benchmarking Hub.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.