llm leaderboard

Best AI for design

Design here means producing an interface a person prefers, not writing about design.

Aldena runs these models inside your team rooms. See what each one costs.

116 of 116 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarkseloeloeloelo
1
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic82.4$5.00 / $25.004 of 4 benchmarks1692171116791670
2
Qwen3.8 Maxqwen/qwen3.8-max
Qwen80.5$2.00 / $6.004 of 4 benchmarks1667167516271631
3
Claude Fable 5anthropic/claude-fable-5
Anthropic79.8$10.00 / $50.004 of 4 benchmarks1627163016691626
4
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI79.2$3.00 / $15.004 of 4 benchmarks1674169216481570
5
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI78.4$5.00 / $30.004 of 4 benchmarks1622162816221581
6
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic77.6$5.00 / $25.003 of 4 benchmarks166316791653
7
Grok 4.6 (high)x-ai/grok-4.6:high
SpaceXAI76.6$2.00 / $6.003 of 4 benchmarks163116401635
8
Grok 4.5x-ai/grok-4.5
SpaceXAI75.1$2.00 / $6.004 of 4 benchmarks1555155615761580
9
Gemini 3.7 Flash (high)google/gemini-3.7-flash:high
Google74.3$0.375 / $1.8753 of 4 benchmarks158715821621
10
DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high
DeepSeek74.2$1.168 / $2.3363 of 4 benchmarks158415941577
11
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic73.8$5.00 / $25.004 of 4 benchmarks1558155515611565
12
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai72.6$0.462 / $1.4523 of 4 benchmarks158515961538
13
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic72.5$5.00 / $25.003 of 4 benchmarks156415631558
14
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic71.7$5.00 / $25.003 of 4 benchmarks155715581553
15
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek71.6$0.0643 / $0.12853 of 4 benchmarks158115881538
16
Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high
Anthropic70.7$5.00 / $25.003 of 4 benchmarks154515381559
17
Muse Spark 1.1meta/muse-spark-1.1
Meta70.2$1.25 / $4.254 of 4 benchmarks1538152915401544
18
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic69.5$5.00 / $25.004 of 4 benchmarks1537152915501538
19
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic69.3$2.00 / $10.004 of 4 benchmarks1540154515201540
20
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic68.0$5.00 / $25.003 of 4 benchmarks153915421525
21
Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh
Meta67.3$1.25 / $4.253 of 4 benchmarks153515301537
22
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google66.5$0.75 / $3.753 of 4 benchmarks153715291535
23
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic66.2$3.00 / $15.004 of 4 benchmarks1524151315281530
24
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI65.6$1.00 / $6.004 of 4 benchmarks1521151215411525
25
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI64.7$5.00 / $30.004 of 4 benchmarks1507149915431526
26
Qwen3.7 Maxqwen/qwen3.7-max
Qwen64.1$1.475 / $4.4253 of 4 benchmarks151715141520
27
Hy3tencent/hy3
Tencent63.4$0.132 / $0.5283 of 4 benchmarks152215341466
28
GLM 5.1z-ai/glm-5.1
Z.ai62.9$0.966 / $3.0363 of 4 benchmarks151015061519
29
Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview
ByteDance61.64 of 4 benchmarks1522152414931506
30
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI61.6$0.10 / $0.604 of 4 benchmarks1517150515371493
31
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI61.2$0.5415 / $2.284 of 4 benchmarks1509150414941520
32
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI61.0$5.00 / $30.004 of 4 benchmarks1487147815401511
33
Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking
Anthropic60.8$5.00 / $25.001 of 4 benchmarks1579
34
Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k
Anthropic60.4$5.00 / $25.003 of 4 benchmarks149414801512
35
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google58.9$1.50 / $9.004 of 4 benchmarks1499149215141492
36
Qwen3.6 Max Previewqwen/qwen3.6-max-preview
Qwen58.7$1.027 / $6.1623 of 4 benchmarks147914701496
37
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google58.3$1.50 / $9.004 of 4 benchmarks1489148315211490
38
Claude Opus 4.6 (thinking)anthropic/claude-opus-4.6:thinking
Anthropic57.6$5.00 / $25.001 of 4 benchmarks1540
39
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi57.5$0.435 / $0.873 of 4 benchmarks147414621482
40
Claude Opus 4.5anthropic/claude-opus-4.5
Anthropic56.8$5.00 / $25.003 of 4 benchmarks146814571482
41
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI56.7$0.71 / $3.504 of 4 benchmarks1473147314611512
42
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI56.0$2.50 / $15.003 of 4 benchmarks146314531477
43
GPT-5.5openai/gpt-5.5
OpenAI55.8$5.00 / $30.004 of 4 benchmarks1457144515041499
44
MiniMax M3minimax/minimax-m3
MiniMax55.2$0.30 / $1.204 of 4 benchmarks1490148614631475
45
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Google54.3$0.75 / $3.751 of 4 benchmarks1529
46
GPT-5.4 (medium)openai/gpt-5.4:medium
OpenAI53.9$2.50 / $15.003 of 4 benchmarks144214291474
47
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google53.2$2.00 / $12.004 of 4 benchmarks1447143714941481
48
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek52.7$1.168 / $2.3363 of 4 benchmarks144614401436
49
Qwen3.6 Plusqwen/qwen3.6-plus
Qwen52.0$0.325 / $1.954 of 4 benchmarks1460145514611470
50
GLM 4.7z-ai/glm-4.7
Z.ai50.9$0.40 / $1.753 of 4 benchmarks143414291446
51
GLM 5z-ai/glm-5
Z.ai50.4$0.60 / $1.923 of 4 benchmarks143614251454
52
MiMo-V2.5xiaomi/mimo-v2.5
Xiaomi50.0$0.14 / $0.283 of 4 benchmarks143814271427
53
Mimo v2 Proxiaomi/mimo-v2-pro
Xiaomi50.03 of 4 benchmarks143414261439
54
GPT-5 (medium)openai/gpt-5:medium
OpenAI49.5$1.25 / $10.002 of 4 benchmarks14191429
55
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google49.1$0.50 / $3.004 of 4 benchmarks1438142914441455
56
GPT-5.2openai/gpt-5.2
OpenAI49.0$1.75 / $14.002 of 4 benchmarks14181428
57
Gemini 3 Progoogle/gemini-3-pro
Google47.44 of 4 benchmarks1438139614691448
58
Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking
MoonshotAI46.2$0.57 / $2.854 of 4 benchmarks1436142914291441
59
Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b
Qwen45.4$0.39 / $2.343 of 4 benchmarks140013891413
60
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI45.4$1.25 / $10.002 of 4 benchmarks13911401
61
GPT-5.4 Mini (high)openai/gpt-5.4-mini:high
OpenAI45.0$0.75 / $4.503 of 4 benchmarks139713871423
62
GPT-5.3-Codexopenai/gpt-5.3-codex
OpenAI44.6$1.75 / $14.004 of 4 benchmarks1409139814291446
63
Inklingthinkingmachines/inkling
Thinking Machines44.4$0.95 / $4.053 of 4 benchmarks140513981383
64
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google44.4$0.30 / $2.503 of 4 benchmarks144914251426
65
MiniMax M2.7minimax/minimax-m2.7
MiniMax44.3$0.30 / $1.203 of 4 benchmarks139713871401
66
Claude Opus 4.1anthropic/claude-opus-4.1
Anthropic44.0$15.00 / $75.002 of 4 benchmarks13891398
67
Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k
Anthropic43.1$3.00 / $15.003 of 4 benchmarks139213861399
68
MiniMax M2.5minimax/minimax-m2.5
MiniMax42.4$0.22 / $0.903 of 4 benchmarks138413731414
69
GPT-5.4openai/gpt-5.4
OpenAI42.2$2.50 / $15.004 of 4 benchmarks1390138914131449
70
MiniMax M2.1minimax/minimax-m2.1
MiniMax41.5$0.30 / $1.203 of 4 benchmarks138713691401
71
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic41.1$3.00 / $15.003 of 4 benchmarks138613841390
72
Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant
Moonshot AI40.64 of 4 benchmarks1405139114141418
73
Solar Pro 4upstage/solar-pro4
Upstage39.3$0.03 / $0.123 of 4 benchmarks137113571393
74
Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning
xAI38.43 of 4 benchmarks137413671367
75
Gemma 4 31Bgoogle/gemma-4-31b-it
Google38.2$0.10 / $0.343 of 4 benchmarks136413571377
76
Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal
Google37.7$0.50 / $3.004 of 4 benchmarks1383137413981432
77
Gemma 4 26B A4B google/gemma-4-26b-a4b-it
Google37.6$0.12 / $0.403 of 4 benchmarks136213541372
78
Muse Glimmer 30Bmeta/muse-glimmer-30b
Meta37.6$0.35 / $1.503 of 4 benchmarks135913521384
79
GLM 5V Turboz-ai/glm-5v-turbo
Z.ai36.9$1.20 / $4.004 of 4 benchmarks1400138413641423
80
Qwen3.5-27Bqwen/qwen3.5-27b
Qwen36.7$0.195 / $1.563 of 4 benchmarks135813451393
81
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek36.6$0.269 / $0.403 of 4 benchmarks136013501369
82
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI36.4$1.25 / $10.001 of 4 benchmarks1420
83
GLM 4.6z-ai/glm-4.6
Z.ai36.3$0.55 / $2.202 of 4 benchmarks13401350
84
GPT-5.1-Codexopenai/gpt-5.1-codex
OpenAI35.4$1.25 / $10.002 of 4 benchmarks13361346
85
Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b
Qwen34.9$0.29 / $2.403 of 4 benchmarks135813501357
86
Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview
Tencent34.53 of 4 benchmarks135613481360
87
MiniMax M2minimax/minimax-m2
MiniMax32.7$0.255 / $1.022 of 4 benchmarks12971307
88
Laguna M.1poolside/laguna-m.1
Poolside32.33 of 4 benchmarks134713381323
89
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI32.1$1.75 / $14.003 of 4 benchmarks133813301348
90
Mimo v2 Flashxiaomi/mimo-v2-flash
Xiaomi31.43 of 4 benchmarks133013091353
91
Grok 4.3x-ai/grok-4.3
SpaceXAI30.9$1.25 / $2.504 of 4 benchmarks1354135113621375
92
DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp
DeepSeek30.9$0.27 / $0.412 of 4 benchmarks12721282
93
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic30.4$1.00 / $5.003 of 4 benchmarks132613231316
94
Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo
Moonshot AI30.23 of 4 benchmarks132313171327
95
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek30.1$0.269 / $0.403 of 4 benchmarks132513361304
96
KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1
Kwaipilot29.82 of 4 benchmarks12551265
97
GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini
OpenAI28.9$0.25 / $2.002 of 4 benchmarks12441254
98
GPT-5.1openai/gpt-5.1
OpenAI28.2$1.25 / $10.004 of 4 benchmarks1341130513641345
99
Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking
Xiaomi28.03 of 4 benchmarks129212311346
100
Laguna XS.2poolside/laguna-xs.2
Poolside27.63 of 4 benchmarks130312951281
101
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral27.5$1.50 / $7.503 of 4 benchmarks126512551314
102
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen27.23 of 4 benchmarks127312571286
103
Grok 4.1 (thinking)x-ai/grok-4.1:thinking
xAI26.22 of 4 benchmarks12101220
104
Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Qwen26.0$0.225 / $1.803 of 4 benchmarks125012341298
105
Grok Code Fast 1x-ai/grok-code-fast-1
xAI24.62 of 4 benchmarks11641172
106
Qwen3.5-Flashqwen/qwen3.5-flash-02-23
Qwen24.5$0.065 / $0.263 of 4 benchmarks123812241293
107
Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning
xAI24.12 of 4 benchmarks11611170
108
Devstral Medium 2507mistralai/devstral-medium-2507
Mistral AI23.72 of 4 benchmarks10801087
109
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Arcee AI23.6$0.22 / $0.853 of 4 benchmarks123912331241
110
Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning
xAI23.43 of 4 benchmarks124012221252
111
Mistral Large 3mistralai/mistral-large-3
Mistral AI23.33 of 4 benchmarks123012391361
112
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Google21.7$0.25 / $1.504 of 4 benchmarks1254124212801336
113
Gemini 2.5 Progoogle/gemini-2.5-pro
Google21.5$1.25 / $10.003 of 4 benchmarks122612351283
114
Granite 4.1 8Bibm-granite/granite-4.1-8b
IBM21.2$0.05 / $0.103 of 4 benchmarks119211791221
115
Devstral 2mistralai/devstral-2
Mistral AI20.63 of 4 benchmarks119411391216
116
Mercury 2inception/mercury-2
Inception20.2$0.25 / $0.753 of 4 benchmarks116611561199
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 2 of the 4 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.