Best LLM for agents
Agentic tool use covers whether a model drives tools to a finished outcome, stays steerable, and recovers when a command fails. Terminal-Bench and other evaluator-run agentic results identify the model together with its agent harness and effort; they are not pure model-only scores.
Aldena runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | % | % | score | score | score | score | score |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 81.0 | $10.00 / $50.00 | 6 of 8 benchmarks | 80.4% | 12.0 | 8.8 | 10.7 | 1.2 | 14.2 | ||
| 2 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 78.0 | $5.00 / $25.00 | 5 of 8 benchmarks | 12.2 | 11.3 | 15.3 | 1.1 | 13.9 | |||
| 3 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 77.1 | $5.00 / $30.00 | 5 of 8 benchmarks | 10.9 | 9.2 | 10.1 | 1.2 | 10.7 | |||
| 4 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 75.9 | $5.00 / $25.00 | 5 of 8 benchmarks | 11.9 | 7.0 | 18.0 | 1.2 | 14.4 | |||
| 5 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 75.7 | $3.00 / $15.00 | 6 of 8 benchmarks | 82.3% | 10.6 | 7.8 | 15.6 | 1.2 | 8.1 | ||
| 6 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 74.0 | $5.00 / $25.00 | 6 of 8 benchmarks | 75.6% | 6.9 | 8.3 | 5.3 | 1.2 | 11.5 | ||
| 7 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 73.3 | $5.00 / $30.00 | 7 of 8 benchmarks | 83.1% | 75.3% | 8.9 | 9.2 | 5.0 | 1.2 | 14.5 | |
| 8 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 73.3 | $5.00 / $30.00 | 5 of 8 benchmarks | 7.7 | 8.6 | 4.0 | 1.2 | 13.2 | |||
| 9 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 69.4 | $2.00 / $6.00 | 5 of 8 benchmarks | 6.2 | 7.0 | 6.4 | 1.2 | 11.2 | |||
| 10 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 69.1 | $5.00 / $25.00 | 5 of 8 benchmarks | 8.2 | 7.7 | 6.6 | 1.1 | 13.1 | |||
| 11 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 68.8 | $5.00 / $25.00 | 5 of 8 benchmarks | 7.7 | 9.9 | 5.5 | 1.2 | 10.2 | |||
| 12 | GPT-5.5openai/gpt-5.5 | OpenAI | 68.8 | $5.00 / $30.00 | 5 of 8 benchmarks | 6.3 | 7.4 | 3.6 | 1.2 | 11.9 | |||
| 13 | Claude Fable 5 (xhigh)anthropic/claude-fable-5:xhigh | Anthropic | 66.8 | $10.00 / $50.00 | 1 of 8 benchmarks | 83.8% | |||||||
| 14 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 66.1 | $0.462 / $1.452 | 5 of 8 benchmarks | 6.7 | 6.1 | 8.4 | 1.2 | 6.3 | |||
| 15 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 65.7 | $0.50 / $3.00 | 1 of 8 benchmarks | 75.8% | ||||||||
| 16 | MiniMax M2.5 (high)minimax/minimax-m2.5:high | MiniMax | 65.7 | $0.22 / $0.90 | 1 of 8 benchmarks | 75.8% | |||||||
| 17 | Claude Opus 5 (xhigh)anthropic/claude-opus-5:xhigh | Anthropic | 65.6 | $5.00 / $25.00 | 1 of 8 benchmarks | 85.8% | |||||||
| 18 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 64.3 | $2.50 / $15.00 | 5 of 8 benchmarks | 5.0 | 6.0 | 4.3 | 1.2 | 9.6 | |||
| 19 | Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium | Anthropic | 63.8 | $5.00 / $25.00 | 1 of 8 benchmarks | 74.4% | |||||||
| 20 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 63.3 | $10.00 / $50.00 | 1 of 8 benchmarks | 83.3% | |||||||
| 21 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 63.1 | $5.00 / $25.00 | 6 of 8 benchmarks | 78.9% | 9.8 | 8.4 | 9.4 | -0.8 | 9.3 | ||
| 22 | GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | 61.6 | $1.75 / $14.00 | 1 of 8 benchmarks | 72.8% | |||||||
| 23 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 61.6 | $1.75 / $14.00 | 1 of 8 benchmarks | 72.8% | |||||||
| 24 | GLM 5 (high)z-ai/glm-5:high | Z.ai | 61.6 | $0.60 / $1.92 | 1 of 8 benchmarks | 72.8% | |||||||
| 25 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 60.8 | $0.10 / $0.60 | 5 of 8 benchmarks | 4.3 | 1.5 | -1.1 | 1.2 | 11.7 | |||
| 26 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 60.5 | $5.00 / $25.00 | 1 of 8 benchmarks | 82.2% | |||||||
| 27 | Muse Sparkmeta/muse-spark | Meta | 60.5 | — | 1 of 8 benchmarks | 82.2% | |||||||
| 28 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 60.3 | $0.0643 / $0.1285 | 5 of 8 benchmarks | 4.0 | 3.3 | 8.7 | 1.2 | 4.5 | |||
| 29 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 60.2 | $1.25 / $4.25 | 6 of 8 benchmarks | 88.1% | 1.2 | -3.3 | 7.3 | 1.2 | 5.6 | ||
| 30 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 60.2 | $1.00 / $6.00 | 5 of 8 benchmarks | 3.8 | 6.2 | -1.8 | 1.2 | 9.7 | |||
| 31 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 60.2 | $2.00 / $6.00 | 5 of 8 benchmarks | 7.6 | 5.3 | 12.5 | 0.1 | 8.4 | |||
| 32 | Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high | Anthropic | 60.1 | $3.00 / $15.00 | 1 of 8 benchmarks | 71.4% | |||||||
| 33 | Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high | Anthropic | 59.6 | $5.00 / $25.00 | 2 of 8 benchmarks | 76.8% | 69.8% | ||||||
| 34 | Kimi K2.5 (high)moonshotai/kimi-k2.5:high | MoonshotAI | 59.4 | $0.57 / $2.85 | 1 of 8 benchmarks | 70.8% | |||||||
| 35 | GPT-5.6 Solopenai/gpt-5.6-sol | OpenAI | 58.7 | $5.00 / $30.00 | 1 of 8 benchmarks | 81.8% | |||||||
| 36 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 58.6 | $3.00 / $15.00 | 1 of 8 benchmarks | 70.6% | |||||||
| 37 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 58.4 | $2.00 / $6.00 | 1 of 8 benchmarks | 79.3% | |||||||
| 38 | DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high | DeepSeek | 57.9 | $0.269 / $0.40 | 1 of 8 benchmarks | 70.0% | |||||||
| 39 | Gemini 3.7 Flash (high)google/gemini-3.7-flash:high | 57.8 | $0.375 / $1.875 | 5 of 8 benchmarks | 3.6 | 2.8 | 9.8 | 1.2 | 2.3 | ||||
| 40 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 57.6 | — | 2 of 8 benchmarks | 74.2% | 70.3% | |||||||
| 41 | Inkling Smallthinkingmachines/inkling-small | Thinking Machines | 57.6 | $0.45 / $1.20 | 1 of 8 benchmarks | 79.2% | |||||||
| 42 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 57.2 | $5.00 / $25.00 | 5 of 8 benchmarks | 2.1 | 8.3 | 8.8 | -31.0 | 10.9 | |||
| 43 | Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high | 57.1 | — | 1 of 8 benchmarks | 69.6% | ||||||||
| 44 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 57.1 | $2.00 / $10.00 | 6 of 8 benchmarks | 74.6% | 7.1 | 5.2 | 3.1 | 1.1 | 11.2 | ||
| 45 | GPT-5.2openai/gpt-5.2 | OpenAI | 56.4 | $1.75 / $14.00 | 1 of 8 benchmarks | 69.0% | |||||||
| 46 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 55.7 | $15.00 / $75.00 | 1 of 8 benchmarks | 67.6% | |||||||
| 47 | Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high | Anthropic | 54.9 | $1.00 / $5.00 | 1 of 8 benchmarks | 66.6% | |||||||
| 48 | GLM 5.2z-ai/glm-5.2 | Z.ai | 54.1 | $0.462 / $1.452 | 1 of 8 benchmarks | 77.8% | |||||||
| 49 | GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium | OpenAI | 53.8 | $1.25 / $10.00 | 1 of 8 benchmarks | 66.0% | |||||||
| 50 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 53.8 | $1.25 / $10.00 | 1 of 8 benchmarks | 66.0% | |||||||
| 51 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 53.0 | $5.00 / $25.00 | 1 of 8 benchmarks | 76.8% | |||||||
| 52 | GPT-5.6 Terra (max)openai/gpt-5.6-terra:max | OpenAI | 52.9 | $1.00 / $6.00 | 1 of 8 benchmarks | 78.4% | |||||||
| 53 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 52.7 | $0.71 / $3.50 | 5 of 8 benchmarks | 1.1 | -1.8 | 4.4 | 1.2 | -1.2 | |||
| 54 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 52.7 | $1.25 / $10.00 | 1 of 8 benchmarks | 65.0% | |||||||
| 55 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 52.1 | $3.00 / $15.00 | 6 of 8 benchmarks | 69.5% | 3.1 | 2.2 | 0.1 | 1.1 | 11.2 | ||
| 56 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 52.0 | $3.00 / $15.00 | 1 of 8 benchmarks | 64.9% | |||||||
| 57 | Inkling (xhigh)thinkingmachines/inkling:xhigh | Thinking Machines | 51.8 | $0.95 / $4.05 | 1 of 8 benchmarks | 76.0% | |||||||
| 58 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 51.2 | $0.60 / $2.50 | 1 of 8 benchmarks | 63.4% | |||||||
| 59 | MiniMax M2minimax/minimax-m2 | MiniMax | 50.5 | $0.255 / $1.02 | 1 of 8 benchmarks | 61.0% | |||||||
| 60 | Muse Spark 1.1 (xhigh)meta/muse-spark-1.1:xhigh | Meta | 50.1 | $1.25 / $4.25 | 1 of 8 benchmarks | 76.2% | |||||||
| 61 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 49.7 | $0.269 / $0.40 | 1 of 8 benchmarks | 60.0% | |||||||
| 62 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 49.0 | $0.25 / $2.00 | 1 of 8 benchmarks | 59.8% | |||||||
| 63 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 48.4 | $2.50 / $15.00 | 1 of 8 benchmarks | 70.6% | |||||||
| 64 | o3openai/o3 | OpenAI | 48.3 | $2.00 / $8.00 | 1 of 8 benchmarks | 58.4% | |||||||
| 65 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 47.7 | $0.5415 / $2.28 | 5 of 8 benchmarks | -0.6 | -1.1 | 0.6 | 1.2 | -5.5 | |||
| 66 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 47.6 | $1.50 / $9.00 | 6 of 8 benchmarks | 83.6% | -0.3 | -0.5 | 0.6 | 0.3 | -1.7 | |||
| 67 | Devstral Small 2512mistralai/devstral-small-2512 | Mistral AI | 47.5 | — | 1 of 8 benchmarks | 56.4% | |||||||
| 68 | GPT-5.6 Luna (max)openai/gpt-5.6-luna:max | OpenAI | 47.3 | $0.10 / $0.60 | 1 of 8 benchmarks | 75.7% | |||||||
| 69 | GPT-5 Miniopenai/gpt-5-mini | OpenAI | 46.8 | $0.25 / $2.00 | 1 of 8 benchmarks | 56.2% | |||||||
| 70 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 46.5 | $5.00 / $25.00 | 2 of 8 benchmarks | 68.9% | 79.1% | ||||||
| 71 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 45.7 | — | 1 of 8 benchmarks | 55.4% | |||||||
| 72 | GLM 4.6z-ai/glm-4.6 | Z.ai | 45.7 | $0.55 / $2.20 | 1 of 8 benchmarks | 55.4% | |||||||
| 73 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 45.6 | $1.475 / $4.425 | 5 of 8 benchmarks | 0.2 | -0.0 | -0.3 | 0.6 | 5.0 | |||
| 74 | GLM 4.5z-ai/glm-4.5 | Z.ai | 44.5 | $0.60 / $2.20 | 1 of 8 benchmarks | 54.2% | |||||||
| 75 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 44.4 | $2.00 / $12.00 | 5 of 8 benchmarks | -0.4 | 3.2 | 1.5 | 0.9 | -11.5 | ||||
| 76 | GLM 5.1z-ai/glm-5.1 | Z.ai | 44.0 | $0.966 / $3.036 | 6 of 8 benchmarks | 75.6% | 0.6 | 1.7 | 1.8 | -0.5 | -0.9 | ||
| 77 | Devstral 2mistralai/devstral-2 | Mistral AI | 43.8 | — | 1 of 8 benchmarks | 53.8% | |||||||
| 78 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 43.8 | $1.75 / $14.00 | 1 of 8 benchmarks | 67.6% | |||||||
| 79 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 43.5 | $2.00 / $12.00 | 2 of 8 benchmarks | 65.8% | 78.2% | |||||||
| 80 | Gemini 2.5 Progoogle/gemini-2.5-pro | 43.1 | $1.25 / $10.00 | 1 of 8 benchmarks | 53.6% | ||||||||
| 81 | Kimi K2p5moonshotai/kimi-k2p5 | Moonshot AI | 42.6 | — | 1 of 8 benchmarks | 64.4% | |||||||
| 82 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 42.3 | — | 1 of 8 benchmarks | 52.8% | |||||||
| 83 | Gemini 3 Pro (high)google/gemini-3-pro:high | 41.8 | — | 1 of 8 benchmarks | 73.9% | ||||||||
| 84 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 41.7 | $1.168 / $2.336 | 5 of 8 benchmarks | 0.1 | 0.6 | -2.2 | 0.3 | 4.4 | |||
| 85 | o4 Miniopenai/o4-mini | OpenAI | 41.6 | $1.10 / $4.40 | 1 of 8 benchmarks | 45.0% | |||||||
| 86 | Kimi K2 Instructmoonshotai/kimi-k2-instruct | Moonshot AI | 40.8 | — | 1 of 8 benchmarks | 43.8% | |||||||
| 87 | Claude Sonnet 4.5 (thinking)anthropic/claude-sonnet-4.5:thinking | Anthropic | 40.3 | $3.00 / $15.00 | 1 of 8 benchmarks | 59.5% | |||||||
| 88 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 40.2 | $0.75 / $3.75 | 5 of 8 benchmarks | -2.4 | -4.4 | -0.9 | 1.2 | -3.6 | ||||
| 89 | GPT-4.1openai/gpt-4.1 | OpenAI | 40.1 | $2.00 / $8.00 | 1 of 8 benchmarks | 39.6% | |||||||
| 90 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 39.4 | $0.05 / $0.40 | 1 of 8 benchmarks | 34.8% | |||||||
| 91 | GLM 4.7z-ai/glm-4.7 | Z.ai | 39.2 | $0.40 / $1.75 | 1 of 8 benchmarks | 58.1% | |||||||
| 92 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 38.6 | $0.30 / $2.50 | 1 of 8 benchmarks | 28.7% | ||||||||
| 93 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 38.4 | $0.32 / $1.28 | 5 of 8 benchmarks | -2.0 | -5.6 | -1.1 | 0.2 | 6.3 | |||
| 94 | Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high | 38.0 | $0.25 / $1.50 | 1 of 8 benchmarks | 57.1% | ||||||||
| 95 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 37.9 | $0.03 / $0.17 | 1 of 8 benchmarks | 26.0% | |||||||
| 96 | MiniMax M3minimax/minimax-m3 | MiniMax | 37.5 | $0.30 / $1.20 | 5 of 8 benchmarks | -2.5 | -5.1 | -5.8 | 0.7 | 5.9 | |||
| 97 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 37.1 | $0.40 / $1.60 | 1 of 8 benchmarks | 23.9% | |||||||
| 98 | GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh | OpenAI | 36.9 | $0.75 / $4.50 | 1 of 8 benchmarks | 56.7% | |||||||
| 99 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 36.4 | $2.50 / $10.00 | 1 of 8 benchmarks | 21.6% | |||||||
| 100 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 35.7 | $1.25 / $10.00 | 1 of 8 benchmarks | 50.1% | |||||||
| 101 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 35.7 | — | 1 of 8 benchmarks | 21.0% | |||||||
| 102 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 35.5 | $0.0643 / $0.1285 | 5 of 8 benchmarks | -2.2 | -1.0 | -2.0 | -1.3 | 2.6 | |||
| 103 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 35.5 | $0.435 / $0.87 | 5 of 8 benchmarks | -2.0 | -2.0 | -2.8 | 0.0 | 1.6 | |||
| 104 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 34.9 | — | 1 of 8 benchmarks | 13.5% | ||||||||
| 105 | o3 Proopenai/o3-pro | OpenAI | 34.6 | $20.00 / $80.00 | 1 of 8 benchmarks | 44.5% | |||||||
| 106 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 34.2 | — | 1 of 8 benchmarks | 9.1% | |||||||
| 107 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 33.4 | $1.00 / $5.00 | 1 of 8 benchmarks | 40.2% | |||||||
| 108 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 33.4 | — | 1 of 8 benchmarks | 9.0% | |||||||
| 109 | GLM 5.1 (max)z-ai/glm-5.1:max | Z.ai | 33.4 | $0.966 / $3.036 | 1 of 8 benchmarks | 58.7% | |||||||
| 110 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 33.2 | $1.50 / $9.00 | 5 of 8 benchmarks | -3.5 | -3.5 | -8.3 | 0.5 | -0.3 | ||||
| 111 | Hy3tencent/hy3 | Tencent | 32.5 | $0.132 / $0.528 | 5 of 8 benchmarks | -1.3 | -8.6 | -2.5 | -1.3 | 4.0 | |||
| 112 | Inklingthinkingmachines/inkling | Thinking Machines | 32.2 | $0.95 / $4.05 | 5 of 8 benchmarks | -6.6 | -11.8 | -11.7 | 0.6 | 6.8 | |||
| 113 | Grok 4.3 (high)x-ai/grok-4.3:high | SpaceXAI | 30.9 | $1.25 / $2.50 | 5 of 8 benchmarks | -8.5 | -7.2 | -9.7 | 1.1 | -13.3 | |||
| 114 | Grok Build 0.1x-ai/grok-build-0.1 | SpaceXAI | 28.9 | $1.00 / $2.00 | 5 of 8 benchmarks | -9.1 | -8.6 | -5.4 | 1.0 | -21.4 | |||
| 115 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 27.4 | $1.25 / $2.50 | 5 of 8 benchmarks | -14.7 | -5.7 | -10.8 | 1.1 | -41.8 | |||
| 116 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 26.2 | $0.50 / $3.00 | 6 of 8 benchmarks | 62.0% | -8.5 | -3.4 | -7.2 | -0.7 | -21.0 | |||
| 117 | Solar Pro 4upstage/solar-pro4 | Upstage | 24.9 | $0.03 / $0.12 | 5 of 8 benchmarks | -10.1 | -13.3 | -8.4 | 0.5 | -13.6 | |||
| 118 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 24.4 | $0.30 / $1.20 | 5 of 8 benchmarks | -11.2 | -13.7 | -10.9 | 1.0 | -16.8 | |||
| 119 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 23.8 | $1.50 / $7.50 | 5 of 8 benchmarks | -7.0 | -11.0 | -9.7 | -2.9 | -1.7 | |||
| 120 | Gemma 4 31Bgoogle/gemma-4-31b-it | 22.7 | $0.10 / $0.34 | 5 of 8 benchmarks | -18.9 | -9.4 | 0.3 | -31.2 | -51.5 | ||||
| 121 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 21.5 | $0.30 / $2.50 | 5 of 8 benchmarks | -10.4 | -10.3 | -12.7 | -0.3 | -15.1 | ||||
| 122 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 20.3 | $0.60 / $3.60 | 5 of 8 benchmarks | -14.6 | -19.9 | -15.4 | 0.5 | -24.1 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 8 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.
- MCP AtlasMIT
All-1,000-task Pass Rate from Scale's current Performance Comparison, using the standardized MCP loop and 100-tool-call budget.
- SWE-bench
Resolve rates published by the SWE-bench maintainers.
- Terminal-Bench 2.1Apache-2.0
Published Accuracy over 89 Terminal-Bench 2.1 tasks, run and verified by the benchmark team; results are model plus agent harness plus effort, not model-only evaluations.