llm leaderboard

Best LLM for agents

Agentic tool use covers whether a model drives tools to a finished outcome, stays steerable, and recovers when a command fails. Terminal-Bench and other evaluator-run agentic results identify the model together with its agent harness and effort; they are not pure model-only scores.

Aldena runs these models inside your team rooms. See what each one costs.

122 of 122 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarks%%%scorescorescorescorescore
1
Claude Fable 5 (high)anthropic/claude-fable-5:high
Anthropic81.0$10.00 / $50.006 of 8 benchmarks80.4%12.08.810.71.214.2
2
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic78.0$5.00 / $25.005 of 8 benchmarks12.211.315.31.113.9
3
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI77.1$5.00 / $30.005 of 8 benchmarks10.99.210.11.210.7
4
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic75.9$5.00 / $25.005 of 8 benchmarks11.97.018.01.214.4
5
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI75.7$3.00 / $15.006 of 8 benchmarks82.3%10.67.815.61.28.1
6
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic74.0$5.00 / $25.006 of 8 benchmarks75.6%6.98.35.31.211.5
7
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI73.3$5.00 / $30.007 of 8 benchmarks83.1%75.3%8.99.25.01.214.5
8
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI73.3$5.00 / $30.005 of 8 benchmarks7.78.64.01.213.2
9
Grok 4.5x-ai/grok-4.5
SpaceXAI69.4$2.00 / $6.005 of 8 benchmarks6.27.06.41.211.2
10
Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high
Anthropic69.1$5.00 / $25.005 of 8 benchmarks8.27.76.61.113.1
11
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic68.8$5.00 / $25.005 of 8 benchmarks7.79.95.51.210.2
12
GPT-5.5openai/gpt-5.5
OpenAI68.8$5.00 / $30.005 of 8 benchmarks6.37.43.61.211.9
13
Claude Fable 5 (xhigh)anthropic/claude-fable-5:xhigh
Anthropic66.8$10.00 / $50.001 of 8 benchmarks83.8%
14
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai66.1$0.462 / $1.4525 of 8 benchmarks6.76.18.41.26.3
15
Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high
Google65.7$0.50 / $3.001 of 8 benchmarks75.8%
16
MiniMax M2.5 (high)minimax/minimax-m2.5:high
MiniMax65.7$0.22 / $0.901 of 8 benchmarks75.8%
17
Claude Opus 5 (xhigh)anthropic/claude-opus-5:xhigh
Anthropic65.6$5.00 / $25.001 of 8 benchmarks85.8%
18
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI64.3$2.50 / $15.005 of 8 benchmarks5.06.04.31.29.6
19
Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium
Anthropic63.8$5.00 / $25.001 of 8 benchmarks74.4%
20
Claude Fable 5anthropic/claude-fable-5
Anthropic63.3$10.00 / $50.001 of 8 benchmarks83.3%
21
Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high
Anthropic63.1$5.00 / $25.006 of 8 benchmarks78.9%9.88.49.4-0.89.3
22
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI61.6$1.75 / $14.001 of 8 benchmarks72.8%
23
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI61.6$1.75 / $14.001 of 8 benchmarks72.8%
24
GLM 5 (high)z-ai/glm-5:high
Z.ai61.6$0.60 / $1.921 of 8 benchmarks72.8%
25
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI60.8$0.10 / $0.605 of 8 benchmarks4.31.5-1.11.211.7
26
Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max
Anthropic60.5$5.00 / $25.001 of 8 benchmarks82.2%
27
Muse Sparkmeta/muse-spark
Meta60.51 of 8 benchmarks82.2%
28
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek60.3$0.0643 / $0.12855 of 8 benchmarks4.03.38.71.24.5
29
Muse Spark 1.1meta/muse-spark-1.1
Meta60.2$1.25 / $4.256 of 8 benchmarks88.1%1.2-3.37.31.25.6
30
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI60.2$1.00 / $6.005 of 8 benchmarks3.86.2-1.81.29.7
31
Qwen3.8 Maxqwen/qwen3.8-max
Qwen60.2$2.00 / $6.005 of 8 benchmarks7.65.312.50.18.4
32
Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high
Anthropic60.1$3.00 / $15.001 of 8 benchmarks71.4%
33
Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high
Anthropic59.6$5.00 / $25.002 of 8 benchmarks76.8%69.8%
34
Kimi K2.5 (high)moonshotai/kimi-k2.5:high
MoonshotAI59.4$0.57 / $2.851 of 8 benchmarks70.8%
35
GPT-5.6 Solopenai/gpt-5.6-sol
OpenAI58.7$5.00 / $30.001 of 8 benchmarks81.8%
36
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic58.6$3.00 / $15.001 of 8 benchmarks70.6%
37
Grok 4.5 (high)x-ai/grok-4.5:high
SpaceXAI58.4$2.00 / $6.001 of 8 benchmarks79.3%
38
DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high
DeepSeek57.9$0.269 / $0.401 of 8 benchmarks70.0%
39
Gemini 3.7 Flash (high)google/gemini-3.7-flash:high
Google57.8$0.375 / $1.8755 of 8 benchmarks3.62.89.81.22.3
40
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Google57.62 of 8 benchmarks74.2%70.3%
41
Inkling Smallthinkingmachines/inkling-small
Thinking Machines57.6$0.45 / $1.201 of 8 benchmarks79.2%
42
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic57.2$5.00 / $25.005 of 8 benchmarks2.18.38.8-31.010.9
43
Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high
Google57.11 of 8 benchmarks69.6%
44
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic57.1$2.00 / $10.006 of 8 benchmarks74.6%7.15.23.11.111.2
45
GPT-5.2openai/gpt-5.2
OpenAI56.4$1.75 / $14.001 of 8 benchmarks69.0%
46
Claude Opus 4anthropic/claude-opus-4
Anthropic55.7$15.00 / $75.001 of 8 benchmarks67.6%
47
Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high
Anthropic54.9$1.00 / $5.001 of 8 benchmarks66.6%
48
GLM 5.2z-ai/glm-5.2
Z.ai54.1$0.462 / $1.4521 of 8 benchmarks77.8%
49
GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium
OpenAI53.8$1.25 / $10.001 of 8 benchmarks66.0%
50
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI53.8$1.25 / $10.001 of 8 benchmarks66.0%
51
Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max
Anthropic53.0$5.00 / $25.001 of 8 benchmarks76.8%
52
GPT-5.6 Terra (max)openai/gpt-5.6-terra:max
OpenAI52.9$1.00 / $6.001 of 8 benchmarks78.4%
53
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI52.7$0.71 / $3.505 of 8 benchmarks1.1-1.84.41.2-1.2
54
GPT-5 (medium)openai/gpt-5:medium
OpenAI52.7$1.25 / $10.001 of 8 benchmarks65.0%
55
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic52.1$3.00 / $15.006 of 8 benchmarks69.5%3.12.20.11.111.2
56
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic52.0$3.00 / $15.001 of 8 benchmarks64.9%
57
Inkling (xhigh)thinkingmachines/inkling:xhigh
Thinking Machines51.8$0.95 / $4.051 of 8 benchmarks76.0%
58
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
MoonshotAI51.2$0.60 / $2.501 of 8 benchmarks63.4%
59
MiniMax M2minimax/minimax-m2
MiniMax50.5$0.255 / $1.021 of 8 benchmarks61.0%
60
Muse Spark 1.1 (xhigh)meta/muse-spark-1.1:xhigh
Meta50.1$1.25 / $4.251 of 8 benchmarks76.2%
61
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek49.7$0.269 / $0.401 of 8 benchmarks60.0%
62
GPT-5 Mini (medium)openai/gpt-5-mini:medium
OpenAI49.0$0.25 / $2.001 of 8 benchmarks59.8%
63
GPT-5.4 (xhigh)openai/gpt-5.4:xhigh
OpenAI48.4$2.50 / $15.001 of 8 benchmarks70.6%
64
o3openai/o3
OpenAI48.3$2.00 / $8.001 of 8 benchmarks58.4%
65
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI47.7$0.5415 / $2.285 of 8 benchmarks-0.6-1.10.61.2-5.5
66
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google47.6$1.50 / $9.006 of 8 benchmarks83.6%-0.3-0.50.60.3-1.7
67
Devstral Small 2512mistralai/devstral-small-2512
Mistral AI47.51 of 8 benchmarks56.4%
68
GPT-5.6 Luna (max)openai/gpt-5.6-luna:max
OpenAI47.3$0.10 / $0.601 of 8 benchmarks75.7%
69
GPT-5 Miniopenai/gpt-5-mini
OpenAI46.8$0.25 / $2.001 of 8 benchmarks56.2%
70
Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max
Anthropic46.5$5.00 / $25.002 of 8 benchmarks68.9%79.1%
71
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen45.71 of 8 benchmarks55.4%
72
GLM 4.6z-ai/glm-4.6
Z.ai45.7$0.55 / $2.201 of 8 benchmarks55.4%
73
Qwen3.7 Maxqwen/qwen3.7-max
Qwen45.6$1.475 / $4.4255 of 8 benchmarks0.2-0.0-0.30.65.0
74
GLM 4.5z-ai/glm-4.5
Z.ai44.5$0.60 / $2.201 of 8 benchmarks54.2%
75
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google44.4$2.00 / $12.005 of 8 benchmarks-0.43.21.50.9-11.5
76
GLM 5.1z-ai/glm-5.1
Z.ai44.0$0.966 / $3.0366 of 8 benchmarks75.6%0.61.71.8-0.5-0.9
77
Devstral 2mistralai/devstral-2
Mistral AI43.81 of 8 benchmarks53.8%
78
GPT-5.2 (xhigh)openai/gpt-5.2:xhigh
OpenAI43.8$1.75 / $14.001 of 8 benchmarks67.6%
79
Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high
Google43.5$2.00 / $12.002 of 8 benchmarks65.8%78.2%
80
Gemini 2.5 Progoogle/gemini-2.5-pro
Google43.1$1.25 / $10.001 of 8 benchmarks53.6%
81
Kimi K2p5moonshotai/kimi-k2p5
Moonshot AI42.61 of 8 benchmarks64.4%
82
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic42.31 of 8 benchmarks52.8%
83
Gemini 3 Pro (high)google/gemini-3-pro:high
Google41.81 of 8 benchmarks73.9%
84
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek41.7$1.168 / $2.3365 of 8 benchmarks0.10.6-2.20.34.4
85
o4 Miniopenai/o4-mini
OpenAI41.6$1.10 / $4.401 of 8 benchmarks45.0%
86
Kimi K2 Instructmoonshotai/kimi-k2-instruct
Moonshot AI40.81 of 8 benchmarks43.8%
87
Claude Sonnet 4.5 (thinking)anthropic/claude-sonnet-4.5:thinking
Anthropic40.3$3.00 / $15.001 of 8 benchmarks59.5%
88
Gemini 3.6 Flash (high)google/gemini-3.6-flash:high
Google40.2$0.75 / $3.755 of 8 benchmarks-2.4-4.4-0.91.2-3.6
89
GPT-4.1openai/gpt-4.1
OpenAI40.1$2.00 / $8.001 of 8 benchmarks39.6%
90
GPT-5 Nano (medium)openai/gpt-5-nano:medium
OpenAI39.4$0.05 / $0.401 of 8 benchmarks34.8%
91
GLM 4.7z-ai/glm-4.7
Z.ai39.2$0.40 / $1.751 of 8 benchmarks58.1%
92
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google38.6$0.30 / $2.501 of 8 benchmarks28.7%
93
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen38.4$0.32 / $1.285 of 8 benchmarks-2.0-5.6-1.10.26.3
94
Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high
Google38.0$0.25 / $1.501 of 8 benchmarks57.1%
95
gpt-oss-120bopenai/gpt-oss-120b
OpenAI37.9$0.03 / $0.171 of 8 benchmarks26.0%
96
MiniMax M3minimax/minimax-m3
MiniMax37.5$0.30 / $1.205 of 8 benchmarks-2.5-5.1-5.80.75.9
97
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI37.1$0.40 / $1.601 of 8 benchmarks23.9%
98
GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh
OpenAI36.9$0.75 / $4.501 of 8 benchmarks56.7%
99
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20
OpenAI36.4$2.50 / $10.001 of 8 benchmarks21.6%
100
GPT-5.1 (high)openai/gpt-5.1:high
OpenAI35.7$1.25 / $10.001 of 8 benchmarks50.1%
101
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta35.71 of 8 benchmarks21.0%
102
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash
DeepSeek35.5$0.0643 / $0.12855 of 8 benchmarks-2.2-1.0-2.0-1.32.6
103
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi35.5$0.435 / $0.875 of 8 benchmarks-2.0-2.0-2.80.01.6
104
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google34.91 of 8 benchmarks13.5%
105
o3 Proopenai/o3-pro
OpenAI34.6$20.00 / $80.001 of 8 benchmarks44.5%
106
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta34.21 of 8 benchmarks9.1%
107
Claude Haiku 4.5anthropic/claude-haiku-4.5
Anthropic33.4$1.00 / $5.001 of 8 benchmarks40.2%
108
Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct
Qwen33.41 of 8 benchmarks9.0%
109
GLM 5.1 (max)z-ai/glm-5.1:max
Z.ai33.4$0.966 / $3.0361 of 8 benchmarks58.7%
110
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google33.2$1.50 / $9.005 of 8 benchmarks-3.5-3.5-8.30.5-0.3
111
Hy3tencent/hy3
Tencent32.5$0.132 / $0.5285 of 8 benchmarks-1.3-8.6-2.5-1.34.0
112
Inklingthinkingmachines/inkling
Thinking Machines32.2$0.95 / $4.055 of 8 benchmarks-6.6-11.8-11.70.66.8
113
Grok 4.3 (high)x-ai/grok-4.3:high
SpaceXAI30.9$1.25 / $2.505 of 8 benchmarks-8.5-7.2-9.71.1-13.3
114
Grok Build 0.1x-ai/grok-build-0.1
SpaceXAI28.9$1.00 / $2.005 of 8 benchmarks-9.1-8.6-5.41.0-21.4
115
Grok 4.3x-ai/grok-4.3
SpaceXAI27.4$1.25 / $2.505 of 8 benchmarks-14.7-5.7-10.81.1-41.8
116
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google26.2$0.50 / $3.006 of 8 benchmarks62.0%-8.5-3.4-7.2-0.7-21.0
117
Solar Pro 4upstage/solar-pro4
Upstage24.9$0.03 / $0.125 of 8 benchmarks-10.1-13.3-8.40.5-13.6
118
MiniMax M2.7minimax/minimax-m2.7
MiniMax24.4$0.30 / $1.205 of 8 benchmarks-11.2-13.7-10.91.0-16.8
119
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral23.8$1.50 / $7.505 of 8 benchmarks-7.0-11.0-9.7-2.9-1.7
120
Gemma 4 31Bgoogle/gemma-4-31b-it
Google22.7$0.10 / $0.345 of 8 benchmarks-18.9-9.40.3-31.2-51.5
121
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google21.5$0.30 / $2.505 of 8 benchmarks-10.4-10.3-12.7-0.3-15.1
122
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
NVIDIA20.3$0.60 / $3.605 of 8 benchmarks-14.6-19.9-15.40.5-24.1
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 8 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.

  • MCP AtlasMIT

    All-1,000-task Pass Rate from Scale's current Performance Comparison, using the standardized MCP loop and 100-tool-call budget.

  • SWE-bench

    Resolve rates published by the SWE-bench maintainers.

  • Terminal-Bench 2.1Apache-2.0

    Published Accuracy over 89 Terminal-Bench 2.1 tasks, run and verified by the benchmark team; results are model plus agent harness plus effort, not model-only evaluations.