llm leaderboard

Best LLM for coding

Issue resolution, code generation, private mergeability tasks, cohort-relative engineering performance, webdev preference and security review contribute once each. LiveCodeBench is the stable 2025-08-01 snapshot of 454 problems dated 2024-08-01 through 2025-05-01, so it does not cover newer frontier families. Every row prints coverage and marks partial evidence.

Aldena runs these models inside your team rooms. See what each one costs.

389 of 389 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarks%%%%%%%%%%%eloelo%count
1
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic86.1$5.00 / $25.005 of 15 benchmarks73.7%48.0%50.0%15301692
2
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic82.2$5.00 / $25.004 of 15 benchmarks72.8%48.0%15311663
3
Claude Fable 5anthropic/claude-fable-5
Anthropic81.8$10.00 / $50.003 of 15 benchmarks88.2%15541627
4
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI80.4$5.00 / $30.004 of 15 benchmarks70.7%46.8%15261622
5
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI79.1$3.00 / $15.005 of 15 benchmarks68.5%1543167437.2%44.0
6
Claude Fable 5 (high)anthropic/claude-fable-5:high
Anthropic78.8$10.00 / $50.003 of 15 benchmarks68.6%52.7%63.9%
7
Claude Fable 5 (max)anthropic/claude-fable-5:max
Anthropic78.7$10.00 / $50.003 of 15 benchmarks69.7%51.6%39.5%
8
Claude Opus 4.5anthropic/claude-opus-4.5
Anthropic77.3$5.00 / $25.005 of 15 benchmarks76.7%52.6%70.7%15231468
9
GPT-5.6 Sol (max)openai/gpt-5.6-sol:max
OpenAI77.0$5.00 / $30.003 of 15 benchmarks72.7%47.5%39.0%
10
Qwen3.8 Maxqwen/qwen3.8-max
Qwen76.7$2.00 / $6.002 of 15 benchmarks15291667
11
Claude Fable 5 (xhigh)anthropic/claude-fable-5:xhigh
Anthropic76.3$10.00 / $50.002 of 15 benchmarks69.9%53.5%
12
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic75.8$5.00 / $25.005 of 15 benchmarks78.7%72.0%48.9%15471537
13
Grok 4.6 (high)x-ai/grok-4.6:high
SpaceXAI75.5$2.00 / $6.004 of 15 benchmarks65.2%48.0%15121631
14
Grok 4.5x-ai/grok-4.5
SpaceXAI74.6$2.00 / $6.003 of 15 benchmarks72.1%15211555
15
Claude Opus 5 (medium)anthropic/claude-opus-5:medium
Anthropic74.3$5.00 / $25.002 of 15 benchmarks68.9%53.4%
16
Muse Spark 1.1meta/muse-spark-1.1
Meta73.5$1.25 / $4.252 of 15 benchmarks15311538
17
Claude Opus 5 (xhigh)anthropic/claude-opus-5:xhigh
Anthropic73.0$5.00 / $25.002 of 15 benchmarks73.2%43.6%
18
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google72.2$0.50 / $3.004 of 15 benchmarks75.4%72.7%15081438
19
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic72.0$5.00 / $25.003 of 15 benchmarks66.5%15241539
20
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic71.5$5.00 / $25.003 of 15 benchmarks56.3%15471558
21
o4 Mini Highopenai/o4-mini-high
OpenAI71.2$1.10 / $4.401 of 15 benchmarks80.2%
22
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic71.0$5.00 / $25.004 of 15 benchmarks35.2%31.1%15521557
23
GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh
OpenAI70.81 of 15 benchmarks80.6%
24
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI70.5$5.00 / $30.004 of 15 benchmarks67.0%43.0%34.3%1507
25
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic70.2$3.00 / $15.005 of 15 benchmarks75.2%1528152429.1%32.0
26
Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k
Anthropic70.0$5.00 / $25.002 of 15 benchmarks15301494
27
Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium
Anthropic70.0$5.00 / $25.001 of 15 benchmarks79.2%
28
Qwen3.6 Max Previewqwen/qwen3.6-max-preview
Qwen70.0$1.027 / $6.1623 of 15 benchmarks76.7%15091479
29
Claude Fable 5 (medium)anthropic/claude-fable-5:medium
Anthropic69.9$10.00 / $50.002 of 15 benchmarks65.4%49.8%
30
o3 (high)openai/o3:high
OpenAI69.9$2.00 / $8.001 of 15 benchmarks75.8%
31
Gemini 3.7 Flash (high)google/gemini-3.7-flash:high
Google69.8$0.375 / $1.8753 of 15 benchmarks65.3%42.2%1587
32
Doubao-Seed-Codebytedance/doubao-seed-code
ByteDance69.71 of 15 benchmarks78.8%
33
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI69.2$5.00 / $30.007 of 15 benchmarks64.4%40.6%10.0%1520148747.7%72.0
34
Muse Sparkmeta/muse-spark
Meta69.11 of 15 benchmarks1526
35
GLM-5.3z-ai/glm-5.3
Z.ai69.11 of 15 benchmarks78.1%
36
Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh
Meta68.9$1.25 / $4.253 of 15 benchmarks54.9%15331535
37
Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max
Anthropic68.9$5.00 / $25.002 of 15 benchmarks46.5%28.6%
38
DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high
DeepSeek68.8$1.168 / $2.3362 of 15 benchmarks14891584
39
GPT-5.6 Sol (high)openai/gpt-5.6-sol:high
OpenAI68.6$5.00 / $30.003 of 15 benchmarks69.4%45.1%20.0%
40
o4 Mini (medium)openai/o4-mini:medium
OpenAI68.5$1.10 / $4.401 of 15 benchmarks74.2%
41
Hy3tencent/hy3
Tencent67.8$0.132 / $0.5282 of 15 benchmarks15031522
42
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek67.6$0.0643 / $0.12852 of 15 benchmarks14801581
43
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi67.6$0.435 / $0.872 of 15 benchmarks15201474
44
DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max
DeepSeek67.5$1.168 / $2.3362 of 15 benchmarks77.6%62.8%
45
Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high
Anthropic67.4$5.00 / $25.001 of 15 benchmarks76.8%
46
GPT-5.6 Terra (max)openai/gpt-5.6-terra:max
OpenAI67.2$1.00 / $6.002 of 15 benchmarks69.6%41.3%
47
GPT-5.2 Chatopenai/gpt-5.2-chat
OpenAI67.1$1.75 / $14.001 of 15 benchmarks1515
48
Grok 4.6x-ai/grok-4.6
SpaceXAI67.0$2.00 / $6.001 of 15 benchmarks77.9%
49
Qwen3.7 Maxqwen/qwen3.7-max
Qwen66.8$1.475 / $4.4254 of 15 benchmarks77.3%9.5%15251517
50
Ernie 5.1baidu/ernie-5.1
Baidu66.71 of 15 benchmarks1514
51
GPT 5.5 Instantopenai/gpt-5.5-instant
OpenAI66.71 of 15 benchmarks1514
52
Qwen3.5 Max Previewqwen/qwen3.5-max-preview
Qwen66.41 of 15 benchmarks1513
53
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI66.4$2.50 / $15.004 of 15 benchmarks76.9%15.6%15211463
54
Gemini 3.7 Flash (medium)google/gemini-3.7-flash:medium
Google66.3$0.375 / $1.8752 of 15 benchmarks65.5%43.6%
55
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI66.2$1.00 / $6.004 of 15 benchmarks60.2%38.8%15161521
56
Dola Seed 2.0 Probytedance/dola-seed-2.0-pro
ByteDance66.11 of 15 benchmarks1513
57
Claude Fable 5 (low)anthropic/claude-fable-5:low
Anthropic66.1$10.00 / $50.002 of 15 benchmarks59.6%48.0%
58
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google66.0$1.50 / $9.002 of 15 benchmarks15071489
59
Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k
Anthropic66.0$15.00 / $75.001 of 15 benchmarks1512
60
Grok 4.5 (high)x-ai/grok-4.5:high
SpaceXAI65.7$2.00 / $6.004 of 15 benchmarks53.8%42.4%38.4%41.0
61
Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high
Google65.7$0.50 / $3.001 of 15 benchmarks75.8%
62
MiniMax M2.5 (high)minimax/minimax-m2.5:high
MiniMax65.7$0.22 / $0.901 of 15 benchmarks75.8%
63
Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max
Anthropic65.5$5.00 / $25.003 of 15 benchmarks83.5%38.5%19.1%
64
Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309
xAI65.11 of 15 benchmarks1508
65
Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools
Google65.1$2.00 / $12.001 of 15 benchmarks75.6%
66
Kimi K3 (none)moonshotai/kimi-k3:none
MoonshotAI65.0$3.00 / $15.001 of 15 benchmarks44.2%
67
Gemini 3 Progoogle/gemini-3-pro
Google65.03 of 15 benchmarks68.7%15181438
68
MiniMax M3minimax/minimax-m3
MiniMax64.8$0.30 / $1.202 of 15 benchmarks14971490
69
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen64.7$0.32 / $1.281 of 15 benchmarks1506
70
EXAONE 4.0 32Blg-ai/exaone-4.0-32b
LG AI Research64.51 of 15 benchmarks70.0%
71
Grok 4.6 (medium)x-ai/grok-4.6:medium
SpaceXAI64.4$2.00 / $6.001 of 15 benchmarks67.5%
72
GPT-5.5openai/gpt-5.5
OpenAI64.3$5.00 / $30.003 of 15 benchmarks64.5%15101457
73
GLM 5z-ai/glm-5
Z.ai64.1$0.60 / $1.924 of 15 benchmarks72.1%69.7%14971436
74
GPT-5.3-Codex (high)openai/gpt-5.3-codex:high
OpenAI64.0$1.75 / $14.001 of 15 benchmarks74.8%
75
Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview
ByteDance64.01 of 15 benchmarks1522
76
R1 0528deepseek/deepseek-r1-0528
DeepSeek63.7$0.50 / $2.152 of 15 benchmarks73.1%1464
77
GPT-5openai/gpt-5
OpenAI63.6$1.25 / $10.001 of 15 benchmarks74.4%
78
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic63.5$2.00 / $10.004 of 15 benchmarks48.2%39.4%15231540
79
Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k
Anthropic63.5$15.00 / $75.001 of 15 benchmarks1499
80
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google63.4$0.30 / $2.502 of 15 benchmarks15031449
81
OpenReasoning Nemotron 32Bnvidia/openreasoning-nemotron-32b
NVIDIA63.21 of 15 benchmarks69.8%
82
GPT-5.6 Luna (max)openai/gpt-5.6-luna:max
OpenAI63.0$0.10 / $0.602 of 15 benchmarks67.2%39.8%
83
GLM 5.1z-ai/glm-5.1
Z.ai62.9$0.966 / $3.0364 of 15 benchmarks74.2%25.7%15151510
84
GLM 5.2z-ai/glm-5.2
Z.ai62.9$0.462 / $1.4521 of 15 benchmarks67.5%
85
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI62.8$0.10 / $0.604 of 15 benchmarks56.9%38.9%14991517
86
Grok 4.6 (xhigh)x-ai/grok-4.6:xhigh
SpaceXAI62.7$2.00 / $6.001 of 15 benchmarks66.7%
87
GPT-5.3 Chatopenai/gpt-5.3-chat
OpenAI62.51 of 15 benchmarks1496
88
Claude Opus 4.8 (xhigh)anthropic/claude-opus-4.8:xhigh
Anthropic62.2$5.00 / $25.002 of 15 benchmarks54.4%45.5%
89
Grok 4.1x-ai/grok-4.1
xAI62.11 of 15 benchmarks1492
90
GPT-5.6 Luna (high)openai/gpt-5.6-luna:high
OpenAI62.1$0.10 / $0.604 of 15 benchmarks44.3%35.9%41.9%57.0
91
SWE-1.7 (none)cognition/swe-1.7:none
Cognition61.51 of 15 benchmarks42.0%
92
Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k
Anthropic61.5$3.00 / $15.002 of 15 benchmarks15191392
93
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI61.5$0.57 / $2.852 of 15 benchmarks15021436
94
o3openai/o3
OpenAI61.5$2.00 / $8.003 of 15 benchmarks58.4%36.0%1460
95
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai61.4$0.462 / $1.4525 of 15 benchmarks78.7%43.8%9.5%15061585
96
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Google61.31 of 15 benchmarks72.9%
97
Mimo v2 Proxiaomi/mimo-v2-pro
Xiaomi61.22 of 15 benchmarks15031434
98
Ernie 5.0 0110baidu/ernie-5.0-0110
Baidu61.21 of 15 benchmarks1490
99
GLM 5 (high)z-ai/glm-5:high
Z.ai60.8$0.60 / $1.921 of 15 benchmarks72.8%
100
Mimo v2 Omnixiaomi/mimo-v2-omni
Xiaomi60.81 of 15 benchmarks1486
101
Grok 4.5 (medium)x-ai/grok-4.5:medium
SpaceXAI60.7$2.00 / $6.001 of 15 benchmarks41.9%
102
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi60.6$0.14 / $0.282 of 15 benchmarks14911438
103
OpenCodeReasoning Nemotron 1.1 32Bnvidia/opencodereasoning-nemotron-1.1-32b
NVIDIA60.51 of 15 benchmarks66.8%
104
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI60.42 of 15 benchmarks15051405
105
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash
DeepSeek60.4$0.0643 / $0.12851 of 15 benchmarks1483
106
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic60.35 of 15 benchmarks61.0%51.7%33.8%31.3%1430
107
Claude Opus 5 (low)anthropic/claude-opus-5:low
Anthropic60.2$5.00 / $25.002 of 15 benchmarks58.1%42.0%
108
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI59.9$1.25 / $10.002 of 15 benchmarks68.0%1491
109
GPT-5.6 Sol (medium)openai/gpt-5.6-sol:medium
OpenAI59.5$5.00 / $30.002 of 15 benchmarks61.1%39.9%
110
Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high
Anthropic59.5$3.00 / $15.001 of 15 benchmarks71.4%
111
GPT-5.2 (xhigh)openai/gpt-5.2:xhigh
OpenAI59.4$1.75 / $14.001 of 15 benchmarks23.0%
112
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
NVIDIA59.3$0.60 / $3.601 of 15 benchmarks1475
113
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google59.1$0.75 / $3.754 of 15 benchmarks46.7%33.7%15221537
114
DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking
DeepSeek59.0$0.27 / $0.411 of 15 benchmarks1475
115
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google58.8$2.00 / $12.004 of 15 benchmarks34.4%14.3%15211447
116
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic58.8$3.00 / $15.005 of 15 benchmarks57.0%58.3%35.6%47.1%1449
117
Kimi K2 0905moonshotai/kimi-k2-0905
MoonshotAI58.7$0.60 / $2.502 of 15 benchmarks71.2%1468
118
LongCat Flash Chatmeituan/longcat-flash-chat
Meituan58.71 of 15 benchmarks1474
119
Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max
Anthropic58.7$2.00 / $10.002 of 15 benchmarks53.9%42.4%
120
GLM 4.7z-ai/glm-4.7
Z.ai58.7$0.40 / $1.752 of 15 benchmarks14851434
121
Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k
Anthropic58.6$3.00 / $15.001 of 15 benchmarks1473
122
Qwen3 Maxqwen/qwen3-max
Qwen58.4$0.78 / $3.901 of 15 benchmarks1473
123
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI58.3$1.75 / $14.003 of 15 benchmarks73.8%66.7%1490
124
Kimi K2.5 (high)moonshotai/kimi-k2.5:high
MoonshotAI58.3$0.57 / $2.851 of 15 benchmarks70.8%
125
Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507
Qwen58.3$0.09 / $0.551 of 15 benchmarks1472
126
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI58.32 of 15 benchmarks15111374
127
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen58.2$0.39 / $2.342 of 15 benchmarks14911400
128
Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203
Baidu58.21 of 15 benchmarks1472
129
Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high
Anthropic57.9$5.00 / $25.005 of 15 benchmarks26.6%1552154526.7%24.0
130
Chatgpt 4oopenai/chatgpt-4o
OpenAI57.81 of 15 benchmarks1468
131
GPT-5 (medium)openai/gpt-5:medium
OpenAI57.7$1.25 / $10.002 of 15 benchmarks71.5%1419
132
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic57.6$5.00 / $25.006 of 15 benchmarks51.8%41.0%1533156424.4%24.0
133
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai57.6$1.20 / $4.002 of 15 benchmarks14901400
134
DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high
DeepSeek57.6$0.269 / $0.401 of 15 benchmarks70.0%
135
GPT-5.4 (medium)openai/gpt-5.4:medium
OpenAI57.4$2.50 / $15.001 of 15 benchmarks1442
136
GPT-5.2openai/gpt-5.2
OpenAI57.4$1.75 / $14.003 of 15 benchmarks69.0%14821418
137
o4 Mini (low)openai/o4-mini:low
OpenAI57.2$1.10 / $4.401 of 15 benchmarks65.9%
138
Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct
Qwen57.1$0.26 / $1.041 of 15 benchmarks1465
139
Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high
Google57.01 of 15 benchmarks69.6%
140
GPT-5.4openai/gpt-5.4
OpenAI57.0$2.50 / $15.003 of 15 benchmarks46.0%15141390
141
GPT-5 (high)openai/gpt-5:high
OpenAI56.9$1.25 / $10.003 of 15 benchmarks73.5%12.7%1469
142
DeepSeek V3.1 Terminus (thinking)deepseek/deepseek-v3.1-terminus:thinking
DeepSeek56.6$0.27 / $0.951 of 15 benchmarks1463
143
GPT 5 Chatopenai/gpt-5-chat
OpenAI56.51 of 15 benchmarks1462
144
Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh
Qwen56.5$2.00 / $6.001 of 15 benchmarks57.5%
145
o3 Mini Highopenai/o3-mini-high
OpenAI56.4$1.10 / $4.402 of 15 benchmarks67.4%1435
146
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google56.1$1.50 / $9.005 of 15 benchmarks79.3%36.1%4.8%15091499
147
MiniMax M2.7minimax/minimax-m2.7
MiniMax56.0$0.30 / $1.202 of 15 benchmarks14791397
148
Gemma 4 31Bgoogle/gemma-4-31b-it
Google56.0$0.10 / $0.342 of 15 benchmarks14991364
149
GPT-5.4 Nano (high)openai/gpt-5.4-nano:high
OpenAI56.0$0.20 / $1.251 of 15 benchmarks1460
150
Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh
Anthropic56.0$2.00 / $10.002 of 15 benchmarks49.7%42.7%
151
Grok 4.5 (low)x-ai/grok-4.5:low
SpaceXAI55.8$2.00 / $6.001 of 15 benchmarks37.9%
152
Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high
Anthropic55.7$1.00 / $5.001 of 15 benchmarks66.6%
153
GPT 4.5 Previewopenai/gpt-4.5-preview
OpenAI55.51 of 15 benchmarks1459
154
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google55.4$0.50 / $3.002 of 15 benchmarks14911383
155
XBai o4 (medium)metastone/xbai-o4:medium
MetaStoneTec55.21 of 15 benchmarks65.0%
156
GPT-5.4 (xhigh)openai/gpt-5.4:xhigh
OpenAI55.1$2.50 / $15.002 of 15 benchmarks51.8%25.4%
157
GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium
OpenAI55.1$1.25 / $10.001 of 15 benchmarks66.0%
158
DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking
DeepSeek55.0$0.25 / $0.951 of 15 benchmarks1457
159
Mistral Medium 2508mistralai/mistral-medium-2508
Mistral AI54.71 of 15 benchmarks1455
160
Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking
Qwen54.6$0.40 / $4.001 of 15 benchmarks1455
161
Kimi K2 0711moonshotai/kimi-k2
MoonshotAI54.6$0.57 / $2.302 of 15 benchmarks65.4%1461
162
Claude Opus 4.5 (128K)anthropic/claude-opus-4.5:128k
Anthropic54.5$5.00 / $25.001 of 15 benchmarks14.3%
163
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic54.5$3.00 / $15.006 of 15 benchmarks71.3%44.3%67.0%2.4%15131386
164
GPT-5.3-Codexopenai/gpt-5.3-codex
OpenAI54.4$1.75 / $14.001 of 15 benchmarks1409
165
DeepSeek V4 Pro (xhigh)deepseek/deepseek-v4-pro:xhigh
DeepSeek54.4$1.168 / $2.3362 of 15 benchmarks26.7%30.0
166
Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k
Anthropic54.31 of 15 benchmarks1452
167
Step 3.5 Flashstepfun/step-3.5-flash
StepFun54.2$0.10 / $0.301 of 15 benchmarks1451
168
GPT-5 Mini (medium)openai/gpt-5-mini:medium
OpenAI54.1$0.25 / $2.001 of 15 benchmarks64.7%
169
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek53.9$1.168 / $2.3363 of 15 benchmarks24.6%15021446
170
Claude Opus 4.1anthropic/claude-opus-4.1
Anthropic53.9$15.00 / $75.004 of 15 benchmarks73.3%7.9%15051389
171
DeepSeek V3.1deepseek/deepseek-chat-v3.1
DeepSeek53.6$0.25 / $0.951 of 15 benchmarks1448
172
Claude 3.5 Haikuanthropic/claude-3-5-haiku
Anthropic53.52 of 15 benchmarks41.7%1385
173
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
MoonshotAI53.4$0.60 / $2.501 of 15 benchmarks63.4%
174
Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct
Qwen53.4$0.10 / $1.101 of 15 benchmarks1446
175
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google53.3$0.12 / $0.402 of 15 benchmarks14811362
176
Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking
Qwen53.2$0.455 / $1.821 of 15 benchmarks1446
177
GPT-5.5 (medium)openai/gpt-5.5:medium
OpenAI53.2$5.00 / $30.002 of 15 benchmarks54.0%36.6%
178
R1deepseek/deepseek-r1
DeepSeek53.1$0.70 / $2.501 of 15 benchmarks1445
179
Muse Glimmer 30Bmeta/muse-glimmer-30b
Meta52.9$0.35 / $1.502 of 15 benchmarks14811359
180
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI52.9$0.5415 / $2.285 of 15 benchmarks76.7%22.2%2.4%15141509
181
Grok 3 Betax-ai/grok-3-beta
xAI52.81 of 15 benchmarks1443
182
Trinity Large Previewarcee-ai/trinity-large-preview
Arcee AI52.71 of 15 benchmarks1443
183
Grok 4.3x-ai/grok-4.3
SpaceXAI52.7$1.25 / $2.502 of 15 benchmarks14881354
184
o3 (medium)openai/o3:medium
OpenAI52.6$2.00 / $8.001 of 15 benchmarks62.3%
185
Gemini 3.7 Flash (low)google/gemini-3.7-flash:low
Google52.6$0.375 / $1.8752 of 15 benchmarks53.8%36.9%
186
Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507
Qwen52.5$0.23 / $2.301 of 15 benchmarks1442
187
Qwen3 235B A22Bqwen/qwen3-235b-a22b
Qwen52.5$0.455 / $1.822 of 15 benchmarks65.9%1433
188
Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507
Qwen52.3$0.0482 / $0.19311 of 15 benchmarks1440
189
GPT-5.6 Terra (high)openai/gpt-5.6-terra:high
OpenAI52.2$1.00 / $6.002 of 15 benchmarks53.8%36.9%
190
DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus
DeepSeek52.1$0.27 / $0.951 of 15 benchmarks1439
191
Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking
Tencent52.01 of 15 benchmarks1438
192
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI51.9$1.25 / $10.002 of 15 benchmarks66.0%1391
193
Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh
Anthropic51.9$5.00 / $25.001 of 15 benchmarks34.9%
194
Gemini 3.6 Flash (medium)google/gemini-3.6-flash:medium
Google51.4$0.75 / $3.751 of 15 benchmarks34.4%
195
Grok 4 0709x-ai/grok-4-0709
xAI51.31 of 15 benchmarks1435
196
o3 Mini (low)openai/o3-mini:low
OpenAI51.2$1.10 / $4.401 of 15 benchmarks57.0%
197
DeepSeek V4 Flash 0423 (max)deepseek/deepseek-v4-flash:max
DeepSeek51.1$0.0643 / $0.12851 of 15 benchmarks53.3%
198
Muse Spark 1.1 (xhigh)meta/muse-spark-1.1:xhigh
Meta51.1$1.25 / $4.251 of 15 benchmarks53.3%
199
MiniMax M2.5minimax/minimax-m2.5
MiniMax51.0$0.22 / $0.903 of 15 benchmarks68.3%14441384
200
Claude Opus 4anthropic/claude-opus-4
Anthropic50.9$15.00 / $75.003 of 15 benchmarks70.7%46.9%1464
201
Mistral Medium 2505mistralai/mistral-medium-2505
Mistral AI50.91 of 15 benchmarks1433
202
GPT-5.1openai/gpt-5.1
OpenAI50.7$1.25 / $10.002 of 15 benchmarks14741341
203
Gemini 2.5 Progoogle/gemini-2.5-pro
Google50.6$1.25 / $10.004 of 15 benchmarks57.6%73.6%14651226
204
Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max
Anthropic50.6$5.00 / $25.001 of 15 benchmarks12.7%
205
Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high
xAI50.62 of 15 benchmarks66.7%1390
206
Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo
Moonshot AI50.42 of 15 benchmarks14861323
207
Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022
Baidu50.31 of 15 benchmarks1432
208
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek50.1$0.269 / $0.403 of 15 benchmarks60.0%14751360
209
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI50.1$0.75 / $4.503 of 15 benchmarks23.1%14971397
210
Kimi K2.5moonshotai/kimi-k2.5
MoonshotAI50.1$0.57 / $2.853 of 15 benchmarks73.8%67.3%23.4%
211
GPT-5 Mini (high)openai/gpt-5-mini:high
OpenAI50.1$0.25 / $2.001 of 15 benchmarks1431
212
o1openai/o1
OpenAI50.0$15.00 / $60.002 of 15 benchmarks64.6%1433
213
Claude Opus 4 (thinking)anthropic/claude-opus-4:thinking
Anthropic49.9$15.00 / $75.001 of 15 benchmarks56.6%
214
o4 Miniopenai/o4-mini
OpenAI49.8$1.10 / $4.403 of 15 benchmarks45.0%33.9%1433
215
DeepSeek V3 0324deepseek/deepseek-chat-v3-0324
DeepSeek49.8$0.27 / $1.121 of 15 benchmarks1429
216
Kimi K2.7 Code (none)moonshotai/kimi-k2.7-code:none
MoonshotAI49.7$0.71 / $3.501 of 15 benchmarks30.1%
217
GLM 4.5 Airz-ai/glm-4.5-air
Z.ai49.6$0.13 / $0.851 of 15 benchmarks1426
218
Devstral Small 2512mistralai/devstral-small-2512
Mistral AI49.61 of 15 benchmarks56.4%
219
GLM 4.7 Flashz-ai/glm-4.7-flash
Z.ai49.5$0.06 / $0.401 of 15 benchmarks1424
220
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen49.4$0.29 / $2.402 of 15 benchmarks14591358
221
Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview
Tencent49.42 of 15 benchmarks14611356
222
Solar Pro 4upstage/solar-pro4
Upstage49.3$0.03 / $0.122 of 15 benchmarks14501371
223
GPT-5.5 (low)openai/gpt-5.5:low
OpenAI49.3$5.00 / $30.004 of 15 benchmarks27.0%30.6%32.6%38.0
224
Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b
NVIDIA49.21 of 15 benchmarks1422
225
MiniMax M2.1minimax/minimax-m2.1
MiniMax49.2$0.30 / $1.202 of 15 benchmarks14401387
226
Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking
Qwen49.1$0.15 / $1.201 of 15 benchmarks1421
227
GLM 4.6Vz-ai/glm-4.6v
Z.ai49.0$0.30 / $0.901 of 15 benchmarks1417
228
Grok 4.1 (thinking)x-ai/grok-4.1:thinking
xAI48.92 of 15 benchmarks14991210
229
GLM 4.5z-ai/glm-4.5
Z.ai48.8$0.60 / $2.202 of 15 benchmarks54.2%1455
230
Claude 3.5 Sonnetanthropic/claude-3-5-sonnet
Anthropic48.86 of 15 benchmarks62.8%51.3%24.9%25.3%36.4%1435
231
MiniMax M1minimax/minimax-m1
MiniMax48.7$0.55 / $2.201 of 15 benchmarks1416
232
Claude Sonnet 5anthropic/claude-sonnet-5
Anthropic48.6$2.00 / $10.002 of 15 benchmarks25.6%27.0
233
Claude Sonnet 4 (thinking)anthropic/claude-sonnet-4:thinking
Anthropic48.5$3.00 / $15.001 of 15 benchmarks56.0%
234
Mistral Small 2506mistralai/mistral-small-2506
Mistral AI48.41 of 15 benchmarks1412
235
Claude Opus 4.7 (low)anthropic/claude-opus-4.7:low
Anthropic48.4$5.00 / $25.001 of 15 benchmarks27.6%
236
Ling Flash 2.0inclusionai/ling-flash-2.0
inclusionAI48.31 of 15 benchmarks1411
237
Composer 2.5cursor/composer-2.5
Cursor48.31 of 15 benchmarks33.8%
238
Qwen3.6 Plusqwen/qwen3.6-plus
Qwen48.1$0.325 / $1.954 of 15 benchmarks57.9%19.9%14951460
239
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral48.1$1.50 / $7.502 of 15 benchmarks14791265
240
INTELLECT-3prime-intellect/intellect-3
Prime Intellect48.11 of 15 benchmarks1409
241
Step 3stepfun/step-3
StepFun48.01 of 15 benchmarks1408
242
Inklingthinkingmachines/inkling
Thinking Machines48.0$0.95 / $4.053 of 15 benchmarks14.0%14941405
243
GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh
OpenAI47.9$0.75 / $4.501 of 15 benchmarks27.0%
244
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen47.93 of 15 benchmarks69.6%14571273
245
Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b
NVIDIA47.9$0.085 / $0.401 of 15 benchmarks1408
246
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen47.9$0.195 / $1.562 of 15 benchmarks14501358
247
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI47.8$0.71 / $3.502 of 15 benchmarks30.5%1473
248
Qwen3 32Bqwen/qwen3-32b
Qwen47.7$0.08 / $0.281 of 15 benchmarks1407
249
Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct
Qwen47.7$0.07 / $0.281 of 15 benchmarks51.6%
250
GLM 4.5Vz-ai/glm-4.5v
Z.ai47.6$0.60 / $1.801 of 15 benchmarks1405
251
Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA47.51 of 15 benchmarks1404
252
Qwen 2.5 (max)qwen/qwen-2.5:max
Qwen47.31 of 15 benchmarks1403
253
Hunyuan T1tencent/hunyuan-t1
Tencent47.21 of 15 benchmarks1399
254
GPT-4.1openai/gpt-4.1
OpenAI47.2$2.00 / $8.003 of 15 benchmarks48.5%31.1%1456
255
Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking
Google47.0$0.10 / $0.401 of 15 benchmarks1397
256
GPT-5.6 Sol (low)openai/gpt-5.6-sol:low
OpenAI47.0$5.00 / $30.002 of 15 benchmarks45.4%35.4%
257
Devstral Small 2505mistralai/devstral-small-2505
Mistral AI46.91 of 15 benchmarks46.8%
258
Nova 2 Liteamazon/nova-2-lite-v1
Amazon46.9$0.30 / $2.501 of 15 benchmarks1395
259
Laguna M.1poolside/laguna-m.1
Poolside46.91 of 15 benchmarks1347
260
Hunyuan TurboStencent/hunyuan-turbos
Tencent46.81 of 15 benchmarks1394
261
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp
DeepSeek46.7$0.27 / $0.412 of 15 benchmarks14651272
262
Composer 2.5 (none)cursor/composer-2.5:none
Cursor46.61 of 15 benchmarks25.6%
263
Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1
NVIDIA46.51 of 15 benchmarks1391
264
Ring Flash 2.0inclusionai/ring-flash-2.0
inclusionAI46.41 of 15 benchmarks1390
265
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI46.3$1.75 / $14.003 of 15 benchmarks72.8%66.3%1338
266
Kimi K2 Instructmoonshotai/kimi-k2-instruct
Moonshot AI46.21 of 15 benchmarks43.8%
267
GLM 5.2 (none)z-ai/glm-5.2:none
Z.ai46.2$0.462 / $1.4521 of 15 benchmarks24.5%
268
Command A (03-2025)cohere/command-a-03-2025
Cohere46.01 of 15 benchmarks1390
269
o3 Miniopenai/o3-mini
OpenAI45.8$1.10 / $4.404 of 15 benchmarks42.4%32.3%63.0%1416
270
Mimo v2 Flashxiaomi/mimo-v2-flash
Xiaomi45.82 of 15 benchmarks14461330
271
Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max
Anthropic45.8$3.00 / $15.001 of 15 benchmarks24.3%
272
Magistral Medium 2506mistralai/magistral-medium-2506
Mistral AI45.71 of 15 benchmarks1387
273
GPT-5.1-Codexopenai/gpt-5.1-codex
OpenAI45.7$1.25 / $10.001 of 15 benchmarks1336
274
Nova Premier 1.0amazon/nova-premier-v1
Amazon45.6$2.50 / $12.501 of 15 benchmarks42.4%
275
O1 Miniopenai/o1-mini
OpenAI45.51 of 15 benchmarks1387
276
Grok 3 Mini Betax-ai/grok-3-mini-beta
xAI45.41 of 15 benchmarks1386
277
Qwen3 30B A3Bqwen/qwen3-30b-a3b
Qwen45.3$0.12 / $0.501 of 15 benchmarks1386
278
Grok 4.6 (low)x-ai/grok-4.6:low
SpaceXAI45.2$2.00 / $6.001 of 15 benchmarks41.6%
279
Mistral Large 3mistralai/mistral-large-3
Mistral AI45.02 of 15 benchmarks14681230
280
GLM 4.6z-ai/glm-4.6
Z.ai44.9$0.55 / $2.203 of 15 benchmarks55.4%14581340
281
Claude Opus 4.6 (medium)anthropic/claude-opus-4.6:medium
Anthropic44.9$5.00 / $25.001 of 15 benchmarks23.7%
282
QwQ 32Bqwen/qwq-32b
Qwen44.91 of 15 benchmarks1384
283
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI44.82 of 15 benchmarks14611240
284
GPT-5 Nano (high)openai/gpt-5-nano:high
OpenAI44.7$0.05 / $0.401 of 15 benchmarks1384
285
Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking
Google44.6$0.10 / $0.401 of 15 benchmarks1384
286
Claude Sonnet 5 (medium)anthropic/claude-sonnet-5:medium
Anthropic44.6$2.00 / $10.002 of 15 benchmarks39.8%35.2%
287
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google44.5$0.25 / $1.502 of 15 benchmarks14571254
288
OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct
Allen Institute for AI44.41 of 15 benchmarks1382
289
Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low
Anthropic44.3$5.00 / $25.002 of 15 benchmarks40.8%35.1%
290
GPT-4.1 Nanoopenai/gpt-4.1-nano
OpenAI44.3$0.10 / $0.401 of 15 benchmarks1374
291
Laguna XS.2poolside/laguna-xs.2
Poolside44.21 of 15 benchmarks1303
292
gpt-oss-20bopenai/gpt-oss-20b
OpenAI44.0$0.03 / $0.131 of 15 benchmarks1370
293
Devstral Small 2507mistralai/devstral-small-2507
Mistral AI43.91 of 15 benchmarks38.0%
294
DeepSeek V4 Flash 0423 (xhigh)deepseek/deepseek-v4-flash:xhigh
DeepSeek43.8$0.0643 / $0.12852 of 15 benchmarks20.9%27.0
295
Mercuryinception/mercury
Inception43.81 of 15 benchmarks1367
296
Gemini 3.6 Flash (low)google/gemini-3.6-flash:low
Google43.6$0.75 / $3.751 of 15 benchmarks22.8%
297
GPT-5 Nano (medium)openai/gpt-5-nano:medium
OpenAI43.5$0.05 / $0.401 of 15 benchmarks34.8%
298
OLMo 3 32B Thinkallenai/olmo-3-32b-think
Allen Institute for AI43.51 of 15 benchmarks1364
299
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic43.4$1.00 / $5.003 of 15 benchmarks64.7%14791326
300
Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1
NVIDIA43.31 of 15 benchmarks1363
301
Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b
NVIDIA43.2$0.05 / $0.201 of 15 benchmarks1362
302
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20
OpenAI43.2$2.50 / $10.001 of 15 benchmarks31.0%
303
Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium
Anthropic43.1$3.00 / $15.001 of 15 benchmarks21.1%
304
Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503
Mistral AI42.91 of 15 benchmarks1362
305
GPT-5.4 Mini (medium)openai/gpt-5.4-mini:medium
OpenAI42.7$0.75 / $4.501 of 15 benchmarks20.8%
306
Gemma 3 27Bgoogle/gemma-3-27b-it
Google42.7$0.08 / $0.451 of 15 benchmarks1358
307
Gemini 1.5 Pro 002google/gemini-1.5-pro-002
Google42.51 of 15 benchmarks1356
308
Hunyuan Large Visiontencent/hunyuan-large-vision
Tencent42.41 of 15 benchmarks1356
309
KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1
Kwaipilot42.41 of 15 benchmarks1255
310
Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct
Qwen42.3$0.36 / $0.401 of 15 benchmarks1356
311
DeepSeek V4 Flash 0731 (high)deepseek/deepseek-v4-flash-0731:high
DeepSeek42.2$0.14 / $0.281 of 15 benchmarks18.8%
312
Mistral Large 2407mistralai/mistral-large-2407
Mistral42.0$2.00 / $6.001 of 15 benchmarks1354
313
Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking
Xiaomi41.92 of 15 benchmarks14311292
314
GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13
OpenAI41.8$5.00 / $15.004 of 15 benchmarks38.8%31.3%12.0%1369
315
Claude Opus 4.6 (low)anthropic/claude-opus-4.6:low
Anthropic41.8$5.00 / $25.001 of 15 benchmarks18.3%
316
Step 1o Turbo 202506stepfun/step-1o-turbo-202506
StepFun41.71 of 15 benchmarks1352
317
GPT-5.6 Terra (medium)openai/gpt-5.6-terra:medium
OpenAI41.6$1.00 / $6.002 of 15 benchmarks35.1%33.9%
318
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Qwen41.5$0.225 / $1.802 of 15 benchmarks14351250
319
GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini
OpenAI41.5$0.25 / $2.001 of 15 benchmarks1244
320
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google41.4$0.30 / $2.503 of 15 benchmarks28.7%61.9%1424
321
DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none
DeepSeek41.4$1.168 / $2.3361 of 15 benchmarks17.6%
322
Gemini 1.5 Pro 001google/gemini-1.5-pro-001
Google41.31 of 15 benchmarks1347
323
Mistral Large 2411mistralai/mistral-large-2411
Mistral AI41.11 of 15 benchmarks1346
324
DeepSeek V3deepseek/deepseek-chat
DeepSeek41.1$0.2574 / $1.02873 of 15 benchmarks36.7%27.2%1388
325
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI41.1$0.40 / $1.602 of 15 benchmarks23.9%1433
326
Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct
Meta41.0$0.10 / $0.321 of 15 benchmarks1346
327
Qwen3.5-Flashqwen/qwen3.5-flash-02-23
Qwen40.9$0.065 / $0.262 of 15 benchmarks14371238
328
Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0
Amazon40.91 of 15 benchmarks1343
329
Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05
Google40.71 of 15 benchmarks1343
330
Qwen 2.5qwen/qwen-2.5
Qwen40.62 of 15 benchmarks40.2%24.7%
331
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek40.6$0.269 / $0.403 of 15 benchmarks59.0%14701325
332
Claude Opus 4.8 (medium)anthropic/claude-opus-4.8:medium
Anthropic40.6$5.00 / $25.004 of 15 benchmarks48.7%40.5%20.9%19.0
333
Claude Sonnet 4.6 (low)anthropic/claude-sonnet-4.6:low
Anthropic40.5$3.00 / $15.001 of 15 benchmarks15.0%
334
OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think
Allen Institute for AI40.31 of 15 benchmarks1338
335
Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct
Meta40.2$0.40 / $0.401 of 15 benchmarks1333
336
Claude 3 Sonnetanthropic/claude-3-sonnet
Anthropic40.11 of 15 benchmarks1318
337
MiniMax M3 (none)minimax/minimax-m3:none
MiniMax40.1$0.30 / $1.201 of 15 benchmarks14.7%
338
Gemma 3 12Bgoogle/gemma-3-12b-it
Google39.9$0.05 / $0.151 of 15 benchmarks1316
339
Gemini 1.5 Flash 002google/gemini-1.5-flash-002
Google39.81 of 15 benchmarks1316
340
Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501
Mistral AI39.61 of 15 benchmarks1312
341
Gemini 1.5 Flash 001google/gemini-1.5-flash-001
Google39.51 of 15 benchmarks1309
342
Gemma 3n E4Bgoogle/gemma-3n-e4b-it
Google39.41 of 15 benchmarks1308
343
MiniMax M2minimax/minimax-m2
MiniMax39.2$0.255 / $1.023 of 15 benchmarks61.0%13851297
344
Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0
Amazon39.21 of 15 benchmarks1306
345
Qwen3.7 Plus (none)qwen/qwen3.7-plus:none
Qwen39.2$0.32 / $1.281 of 15 benchmarks10.2%
346
Claude Sonnet 5 (low)anthropic/claude-sonnet-5:low
Anthropic39.2$2.00 / $10.002 of 15 benchmarks30.5%28.7%
347
GPT 4 1106 Previewopenai/gpt-4-1106-preview
OpenAI39.14 of 15 benchmarks22.4%28.3%12.5%1340
348
Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning
xAI39.02 of 15 benchmarks14361161
349
Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0
Amazon39.01 of 15 benchmarks1289
350
Command R (08-2024)cohere/command-r-08-2024
Cohere38.8$0.15 / $0.601 of 15 benchmarks1281
351
SWE-1.6 (none)cognition/swe-1.6:none
Cognition38.71 of 15 benchmarks9.4%
352
Command R+ (08-2024)cohere/command-r-plus-08-2024
Cohere38.7$2.50 / $10.001 of 15 benchmarks1280
353
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Arcee AI38.7$0.22 / $0.852 of 15 benchmarks14141239
354
OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct
Allen Institute for AI38.51 of 15 benchmarks1280
355
Grok Code Fast 1x-ai/grok-code-fast-1
xAI38.51 of 15 benchmarks1164
356
Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct
Mistral38.4$2.00 / $6.001 of 15 benchmarks1277
357
GPT-5.4 Mini (low)openai/gpt-5.4-mini:low
OpenAI38.3$0.75 / $4.501 of 15 benchmarks9.3%
358
Gemma 3 4Bgoogle/gemma-3-4b-it
Google38.31 of 15 benchmarks1274
359
Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001
Google38.11 of 15 benchmarks1272
360
Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct
Meta38.0$0.05 / $0.081 of 15 benchmarks1260
361
Devstral Medium 2507mistralai/devstral-medium-2507
Mistral AI37.91 of 15 benchmarks1080
362
Mistral Medium 3.5 (none)mistralai/mistral-medium-3-5:none
Mistral37.9$1.50 / $7.501 of 15 benchmarks8.0%
363
GPT-4oopenai/gpt-4o
OpenAI37.9$2.50 / $10.001 of 15 benchmarks12.2%
364
QwQ 32B Previewqwen/qwq-32b-preview
Qwen37.91 of 15 benchmarks1173
365
gpt-oss-120bopenai/gpt-oss-120b
OpenAI37.8$0.03 / $0.172 of 15 benchmarks26.0%1390
366
Devstral 2mistralai/devstral-2
Mistral AI37.22 of 15 benchmarks53.8%1194
367
GPT-5 Miniopenai/gpt-5-mini
OpenAI36.9$0.25 / $2.002 of 15 benchmarks56.2%39.7%
368
GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06
OpenAI36.9$2.50 / $10.005 of 15 benchmarks27.0%39.7%30.4%29.5%1360
369
GPT-5.6 Luna (medium)openai/gpt-5.6-luna:medium
OpenAI35.7$0.10 / $0.602 of 15 benchmarks11.3%25.7%
370
Mercury 2inception/mercury-2
Inception35.6$0.25 / $0.752 of 15 benchmarks13941166
371
Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high
Anthropic35.4$3.00 / $15.002 of 15 benchmarks29.9%23.5%
372
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta35.42 of 15 benchmarks21.0%1373
373
GPT-5.6 Terra (low)openai/gpt-5.6-terra:low
OpenAI35.3$1.00 / $6.002 of 15 benchmarks24.1%24.1%
374
Claude 3 Opusanthropic/claude-3-opus
Anthropic35.04 of 15 benchmarks15.8%26.3%10.5%1355
375
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google34.42 of 15 benchmarks13.5%1365
376
GPT-4 Turboopenai/gpt-4-turbo
OpenAI34.1$10.00 / $30.002 of 15 benchmarks28.7%1347
377
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta33.72 of 15 benchmarks9.1%1362
378
Claude 2anthropic/claude-2
Anthropic33.73 of 15 benchmarks4.4%3.0%2.0%
379
GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18
OpenAI33.2$0.15 / $0.602 of 15 benchmarks27.5%1349
380
Granite 4.1 8Bibm-granite/granite-4.1-8b
IBM32.3$0.05 / $0.102 of 15 benchmarks13531192
381
Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct
Qwen31.62 of 15 benchmarks9.0%1342
382
GPT-5.6 Luna (low)openai/gpt-5.6-luna:low
OpenAI30.7$0.10 / $0.602 of 15 benchmarks1.6%15.4%
383
Claude Opus 4.7 (medium)anthropic/claude-opus-4.7:medium
Anthropic29.6$5.00 / $25.003 of 15 benchmarks28.9%7.0%7.0
384
Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high
Google29.2$2.00 / $12.002 of 15 benchmarks11.7%8.9%
385
GLM 5.2 (high)z-ai/glm-5.2:high
Z.ai29.2$0.462 / $1.4523 of 15 benchmarks36.3%17.4%18.0
386
SWE Llamaprinceton-nlp/swe-llama
Princeton NLP29.03 of 15 benchmarks1.4%1.3%0.7%
387
Claude 3 Haikuanthropic/claude-3-haiku
Anthropic27.8$0.25 / $1.253 of 15 benchmarks40.6%20.2%1301
388
SWE Llama 13Bprinceton-nlp/swe-llama-13b
Princeton NLP27.43 of 15 benchmarks1.2%1.0%0.7%
389
GPT 3.5openai/gpt-3.5
OpenAI22.73 of 15 benchmarks0.4%0.3%0.2%
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 15 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • DeepSWE v1.1

    Pass@1 across all 113 DeepSWE v1.1 tasks, run by Datacurve with mini-swe-agent through Pier; reasoning-effort configurations remain separate.

  • Epoch AI Benchmarking HubCC BY 4.0

    Benchmark runs by Epoch AI, from the AI Benchmarking Hub.

  • FrontierCode 1.1

    Main weighted rubric Scores on FrontierCode 1.1's 100 hardest tasks, run by Cognition across five trials; results identify model, harness, and effort.

  • FrontierSWE

    Mean@5 Dominance across FrontierSWE's complete 17-task cohort; results identify the evaluated model and harness.

  • LiveCodeBenchMIT

    Maintainer-published Code Generation pass@1 over the 454-problem release_v6 window (2024-08-01 through 2025-05-01), from the stable 2025-08-01 snapshot; it does not cover newer frontier families.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.

  • SWE-bench

    Resolve rates published by the SWE-bench maintainers.

  • Warden

    Security review results published by Warden.