---
title: "Best LLM for agents"
description: "Models ranked on Terminal-Bench 2.1, MCP Atlas, the LMArena agent arena and SWE-bench bash-only, scored on tool use, steerability and recovery."
url: "https://aldena.ai/best-llm-for-agents"
---

# Best LLM for agents

Agentic tool use covers whether a model drives tools to a finished outcome, stays steerable, and recovers when a command fails. Terminal-Bench and other evaluator-run agentic results identify the model together with its agent harness and effort; they are not pure model-only scores.

Aldena runs these models inside your team rooms. [See what each one costs](https://aldena.ai/models).

| rank | model | vendor | score | coverage (benchmarks) | price (in / out) | SWE-bench Bash Only (%) | Terminal-Bench 2.1 (%) | MCP Atlas (%) | LMArena agent (score) | LMArena agent steerability (score) | LMArena agent task outcome (score) | LMArena agent tool hallucination (score) | LMArena agent bash recovery (score) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Fable 5 (high) | Anthropic | 81.0 | 6 of 8 | $10.00 / $50.00 | — | 80.4% | — | 12.0 | 8.8 | 10.7 | 1.2 | 14.2 |
| 2 | Claude Opus 5 (high) | Anthropic | 78.0 | 5 of 8 | $5.00 / $25.00 | — | — | — | 12.2 | 11.3 | 15.3 | 1.1 | 13.9 |
| 3 | GPT-5.6 Sol (xhigh) | OpenAI | 77.1 | 5 of 8 | $5.00 / $30.00 | — | — | — | 10.9 | 9.2 | 10.1 | 1.2 | 10.7 |
| 4 | Claude Opus 5 (max) | Anthropic | 75.9 | 5 of 8 | $5.00 / $25.00 | — | — | — | 11.9 | 7.0 | 18.0 | 1.2 | 14.4 |
| 5 | Kimi K3 (max) | MoonshotAI | 75.7 | 6 of 8 | $3.00 / $15.00 | — | — | 82.3% | 10.6 | 7.8 | 15.6 | 1.2 | 8.1 |
| 6 | Claude Opus 4.6 | Anthropic | 74.0 | 6 of 8 | $5.00 / $25.00 | 75.6% | — | — | 6.9 | 8.3 | 5.3 | 1.2 | 11.5 |
| 7 | GPT-5.5 (xhigh) | OpenAI | 73.3 | 7 of 8 | $5.00 / $30.00 | — | 83.1% | 75.3% | 8.9 | 9.2 | 5.0 | 1.2 | 14.5 |
| 8 | GPT-5.5 (high) | OpenAI | 73.3 | 5 of 8 | $5.00 / $30.00 | — | — | — | 7.7 | 8.6 | 4.0 | 1.2 | 13.2 |
| 9 | Grok 4.5 | SpaceXAI | 69.4 | 5 of 8 | $2.00 / $6.00 | — | — | — | 6.2 | 7.0 | 6.4 | 1.2 | 11.2 |
| 10 | Claude Opus 4.7 (high) | Anthropic | 69.1 | 5 of 8 | $5.00 / $25.00 | — | — | — | 8.2 | 7.7 | 6.6 | 1.1 | 13.1 |
| 11 | Claude Opus 4.7 | Anthropic | 68.8 | 5 of 8 | $5.00 / $25.00 | — | — | — | 7.7 | 9.9 | 5.5 | 1.2 | 10.2 |
| 12 | GPT-5.5 | OpenAI | 68.8 | 5 of 8 | $5.00 / $30.00 | — | — | — | 6.3 | 7.4 | 3.6 | 1.2 | 11.9 |
| 13 | Claude Fable 5 (xhigh) (partial coverage) | Anthropic | 66.8 | 1 of 8 | $10.00 / $50.00 | — | 83.8% | — | — | — | — | — | — |
| 14 | GLM 5.2 (max) | Z.ai | 66.1 | 5 of 8 | $0.462 / $1.452 | — | — | — | 6.7 | 6.1 | 8.4 | 1.2 | 6.3 |
| 15 | Gemini 3 Flash Preview (high) (partial coverage) | Google | 65.7 | 1 of 8 | $0.50 / $3.00 | 75.8% | — | — | — | — | — | — | — |
| 16 | MiniMax M2.5 (high) (partial coverage) | MiniMax | 65.7 | 1 of 8 | $0.22 / $0.90 | 75.8% | — | — | — | — | — | — | — |
| 17 | Claude Opus 5 (xhigh) (partial coverage) | Anthropic | 65.6 | 1 of 8 | $5.00 / $25.00 | — | — | 85.8% | — | — | — | — | — |
| 18 | GPT-5.4 (high) | OpenAI | 64.3 | 5 of 8 | $2.50 / $15.00 | — | — | — | 5.0 | 6.0 | 4.3 | 1.2 | 9.6 |
| 19 | Claude Opus 4.5 (medium) (partial coverage) | Anthropic | 63.8 | 1 of 8 | $5.00 / $25.00 | 74.4% | — | — | — | — | — | — | — |
| 20 | Claude Fable 5 (partial coverage) | Anthropic | 63.3 | 1 of 8 | $10.00 / $50.00 | — | — | 83.3% | — | — | — | — | — |
| 21 | Claude Opus 4.8 (high) | Anthropic | 63.1 | 6 of 8 | $5.00 / $25.00 | — | 78.9% | — | 9.8 | 8.4 | 9.4 | -0.8 | 9.3 |
| 22 | GPT-5.2-Codex (partial coverage) | OpenAI | 61.6 | 1 of 8 | $1.75 / $14.00 | 72.8% | — | — | — | — | — | — | — |
| 23 | GPT-5.2 (high) (partial coverage) | OpenAI | 61.6 | 1 of 8 | $1.75 / $14.00 | 72.8% | — | — | — | — | — | — | — |
| 24 | GLM 5 (high) (partial coverage) | Z.ai | 61.6 | 1 of 8 | $0.60 / $1.92 | 72.8% | — | — | — | — | — | — | — |
| 25 | GPT-5.6 Luna (xhigh) | OpenAI | 60.8 | 5 of 8 | $0.10 / $0.60 | — | — | — | 4.3 | 1.5 | -1.1 | 1.2 | 11.7 |
| 26 | Claude Opus 4.8 (max) (partial coverage) | Anthropic | 60.5 | 1 of 8 | $5.00 / $25.00 | — | — | 82.2% | — | — | — | — | — |
| 27 | Muse Spark (partial coverage) | Meta | 60.5 | 1 of 8 | — | — | — | 82.2% | — | — | — | — | — |
| 28 | DeepSeek V4 Flash 0423 (high) | DeepSeek | 60.3 | 5 of 8 | $0.0643 / $0.1285 | — | — | — | 4.0 | 3.3 | 8.7 | 1.2 | 4.5 |
| 29 | Muse Spark 1.1 | Meta | 60.2 | 6 of 8 | $1.25 / $4.25 | — | — | 88.1% | 1.2 | -3.3 | 7.3 | 1.2 | 5.6 |
| 30 | GPT-5.6 Terra (xhigh) | OpenAI | 60.2 | 5 of 8 | $1.00 / $6.00 | — | — | — | 3.8 | 6.2 | -1.8 | 1.2 | 9.7 |
| 31 | Qwen3.8 Max | Qwen | 60.2 | 5 of 8 | $2.00 / $6.00 | — | — | — | 7.6 | 5.3 | 12.5 | 0.1 | 8.4 |
| 32 | Claude Sonnet 4.5 (high) (partial coverage) | Anthropic | 60.1 | 1 of 8 | $3.00 / $15.00 | 71.4% | — | — | — | — | — | — | — |
| 33 | Claude Opus 4.5 (high) (partial coverage) | Anthropic | 59.6 | 2 of 8 | $5.00 / $25.00 | 76.8% | — | 69.8% | — | — | — | — | — |
| 34 | Kimi K2.5 (high) (partial coverage) | MoonshotAI | 59.4 | 1 of 8 | $0.57 / $2.85 | 70.8% | — | — | — | — | — | — | — |
| 35 | GPT-5.6 Sol (partial coverage) | OpenAI | 58.7 | 1 of 8 | $5.00 / $30.00 | — | — | 81.8% | — | — | — | — | — |
| 36 | Claude Sonnet 4.5 (partial coverage) | Anthropic | 58.6 | 1 of 8 | $3.00 / $15.00 | 70.6% | — | — | — | — | — | — | — |
| 37 | Grok 4.5 (high) (partial coverage) | SpaceXAI | 58.4 | 1 of 8 | $2.00 / $6.00 | — | 79.3% | — | — | — | — | — | — |
| 38 | DeepSeek V3.2 (high) (partial coverage) | DeepSeek | 57.9 | 1 of 8 | $0.269 / $0.40 | 70.0% | — | — | — | — | — | — | — |
| 39 | Gemini 3.7 Flash (high) | Google | 57.8 | 5 of 8 | $0.375 / $1.875 | — | — | — | 3.6 | 2.8 | 9.8 | 1.2 | 2.3 |
| 40 | Gemini 3 Pro Preview (partial coverage) | Google | 57.6 | 2 of 8 | — | 74.2% | — | 70.3% | — | — | — | — | — |
| 41 | Inkling Small (partial coverage) | Thinking Machines | 57.6 | 1 of 8 | $0.45 / $1.20 | — | — | 79.2% | — | — | — | — | — |
| 42 | Claude Opus 4.8 | Anthropic | 57.2 | 5 of 8 | $5.00 / $25.00 | — | — | — | 2.1 | 8.3 | 8.8 | -31.0 | 10.9 |
| 43 | Gemini 3 Pro Preview (high) (partial coverage) | Google | 57.1 | 1 of 8 | — | 69.6% | — | — | — | — | — | — | — |
| 44 | Claude Sonnet 5 (high) | Anthropic | 57.1 | 6 of 8 | $2.00 / $10.00 | — | 74.6% | — | 7.1 | 5.2 | 3.1 | 1.1 | 11.2 |
| 45 | GPT-5.2 (partial coverage) | OpenAI | 56.4 | 1 of 8 | $1.75 / $14.00 | 69.0% | — | — | — | — | — | — | — |
| 46 | Claude Opus 4 (partial coverage) | Anthropic | 55.7 | 1 of 8 | $15.00 / $75.00 | 67.6% | — | — | — | — | — | — | — |
| 47 | Claude Haiku 4.5 (high) (partial coverage) | Anthropic | 54.9 | 1 of 8 | $1.00 / $5.00 | 66.6% | — | — | — | — | — | — | — |
| 48 | GLM 5.2 (partial coverage) | Z.ai | 54.1 | 1 of 8 | $0.462 / $1.452 | — | — | 77.8% | — | — | — | — | — |
| 49 | GPT-5.1-Codex (medium) (partial coverage) | OpenAI | 53.8 | 1 of 8 | $1.25 / $10.00 | 66.0% | — | — | — | — | — | — | — |
| 50 | GPT-5.1 (medium) (partial coverage) | OpenAI | 53.8 | 1 of 8 | $1.25 / $10.00 | 66.0% | — | — | — | — | — | — | — |
| 51 | Claude Opus 4.6 (max) (partial coverage) | Anthropic | 53.0 | 1 of 8 | $5.00 / $25.00 | — | — | 76.8% | — | — | — | — | — |
| 52 | GPT-5.6 Terra (max) (partial coverage) | OpenAI | 52.9 | 1 of 8 | $1.00 / $6.00 | — | 78.4% | — | — | — | — | — | — |
| 53 | Kimi K2.7 Code | MoonshotAI | 52.7 | 5 of 8 | $0.71 / $3.50 | — | — | — | 1.1 | -1.8 | 4.4 | 1.2 | -1.2 |
| 54 | GPT-5 (medium) (partial coverage) | OpenAI | 52.7 | 1 of 8 | $1.25 / $10.00 | 65.0% | — | — | — | — | — | — | — |
| 55 | Claude Sonnet 4.6 | Anthropic | 52.1 | 6 of 8 | $3.00 / $15.00 | — | — | 69.5% | 3.1 | 2.2 | 0.1 | 1.1 | 11.2 |
| 56 | Claude Sonnet 4 (partial coverage) | Anthropic | 52.0 | 1 of 8 | $3.00 / $15.00 | 64.9% | — | — | — | — | — | — | — |
| 57 | Inkling (xhigh) (partial coverage) | Thinking Machines | 51.8 | 1 of 8 | $0.95 / $4.05 | — | — | 76.0% | — | — | — | — | — |
| 58 | Kimi K2 Thinking (partial coverage) | MoonshotAI | 51.2 | 1 of 8 | $0.60 / $2.50 | 63.4% | — | — | — | — | — | — | — |
| 59 | MiniMax M2 (partial coverage) | MiniMax | 50.5 | 1 of 8 | $0.255 / $1.02 | 61.0% | — | — | — | — | — | — | — |
| 60 | Muse Spark 1.1 (xhigh) (partial coverage) | Meta | 50.1 | 1 of 8 | $1.25 / $4.25 | — | 76.2% | — | — | — | — | — | — |
| 61 | DeepSeek V3.2 (thinking) (partial coverage) | DeepSeek | 49.7 | 1 of 8 | $0.269 / $0.40 | 60.0% | — | — | — | — | — | — | — |
| 62 | GPT-5 Mini (medium) (partial coverage) | OpenAI | 49.0 | 1 of 8 | $0.25 / $2.00 | 59.8% | — | — | — | — | — | — | — |
| 63 | GPT-5.4 (xhigh) (partial coverage) | OpenAI | 48.4 | 1 of 8 | $2.50 / $15.00 | — | — | 70.6% | — | — | — | — | — |
| 64 | o3 (partial coverage) | OpenAI | 48.3 | 1 of 8 | $2.00 / $8.00 | 58.4% | — | — | — | — | — | — | — |
| 65 | Kimi K2.6 | MoonshotAI | 47.7 | 5 of 8 | $0.5415 / $2.28 | — | — | — | -0.6 | -1.1 | 0.6 | 1.2 | -5.5 |
| 66 | Gemini 3.5 Flash (high) | Google | 47.6 | 6 of 8 | $1.50 / $9.00 | — | — | 83.6% | -0.3 | -0.5 | 0.6 | 0.3 | -1.7 |
| 67 | Devstral Small 2512 (partial coverage) | Mistral AI | 47.5 | 1 of 8 | — | 56.4% | — | — | — | — | — | — | — |
| 68 | GPT-5.6 Luna (max) (partial coverage) | OpenAI | 47.3 | 1 of 8 | $0.10 / $0.60 | — | 75.7% | — | — | — | — | — | — |
| 69 | GPT-5 Mini (partial coverage) | OpenAI | 46.8 | 1 of 8 | $0.25 / $2.00 | 56.2% | — | — | — | — | — | — | — |
| 70 | Claude Opus 4.7 (max) (partial coverage) | Anthropic | 46.5 | 2 of 8 | $5.00 / $25.00 | — | 68.9% | 79.1% | — | — | — | — | — |
| 71 | Qwen3 Coder 480B A35b Instruct (partial coverage) | Qwen | 45.7 | 1 of 8 | — | 55.4% | — | — | — | — | — | — | — |
| 72 | GLM 4.6 (partial coverage) | Z.ai | 45.7 | 1 of 8 | $0.55 / $2.20 | 55.4% | — | — | — | — | — | — | — |
| 73 | Qwen3.7 Max | Qwen | 45.6 | 5 of 8 | $1.475 / $4.425 | — | — | — | 0.2 | -0.0 | -0.3 | 0.6 | 5.0 |
| 74 | GLM 4.5 (partial coverage) | Z.ai | 44.5 | 1 of 8 | $0.60 / $2.20 | 54.2% | — | — | — | — | — | — | — |
| 75 | Gemini 3.1 Pro Preview | Google | 44.4 | 5 of 8 | $2.00 / $12.00 | — | — | — | -0.4 | 3.2 | 1.5 | 0.9 | -11.5 |
| 76 | GLM 5.1 | Z.ai | 44.0 | 6 of 8 | $0.966 / $3.036 | — | — | 75.6% | 0.6 | 1.7 | 1.8 | -0.5 | -0.9 |
| 77 | Devstral 2 (partial coverage) | Mistral AI | 43.8 | 1 of 8 | — | 53.8% | — | — | — | — | — | — | — |
| 78 | GPT-5.2 (xhigh) (partial coverage) | OpenAI | 43.8 | 1 of 8 | $1.75 / $14.00 | — | — | 67.6% | — | — | — | — | — |
| 79 | Gemini 3.1 Pro Preview (high) (partial coverage) | Google | 43.5 | 2 of 8 | $2.00 / $12.00 | — | 65.8% | 78.2% | — | — | — | — | — |
| 80 | Gemini 2.5 Pro (partial coverage) | Google | 43.1 | 1 of 8 | $1.25 / $10.00 | 53.6% | — | — | — | — | — | — | — |
| 81 | Kimi K2p5 (partial coverage) | Moonshot AI | 42.6 | 1 of 8 | — | — | — | 64.4% | — | — | — | — | — |
| 82 | Claude 3.7 Sonnet (partial coverage) | Anthropic | 42.3 | 1 of 8 | — | 52.8% | — | — | — | — | — | — | — |
| 83 | Gemini 3 Pro (high) (partial coverage) | Google | 41.8 | 1 of 8 | — | — | 73.9% | — | — | — | — | — | — |
| 84 | DeepSeek V4 Pro | DeepSeek | 41.7 | 5 of 8 | $1.168 / $2.336 | — | — | — | 0.1 | 0.6 | -2.2 | 0.3 | 4.4 |
| 85 | o4 Mini (partial coverage) | OpenAI | 41.6 | 1 of 8 | $1.10 / $4.40 | 45.0% | — | — | — | — | — | — | — |
| 86 | Kimi K2 Instruct (partial coverage) | Moonshot AI | 40.8 | 1 of 8 | — | 43.8% | — | — | — | — | — | — | — |
| 87 | Claude Sonnet 4.5 (thinking) (partial coverage) | Anthropic | 40.3 | 1 of 8 | $3.00 / $15.00 | — | — | 59.5% | — | — | — | — | — |
| 88 | Gemini 3.6 Flash (high) | Google | 40.2 | 5 of 8 | $0.75 / $3.75 | — | — | — | -2.4 | -4.4 | -0.9 | 1.2 | -3.6 |
| 89 | GPT-4.1 (partial coverage) | OpenAI | 40.1 | 1 of 8 | $2.00 / $8.00 | 39.6% | — | — | — | — | — | — | — |
| 90 | GPT-5 Nano (medium) (partial coverage) | OpenAI | 39.4 | 1 of 8 | $0.05 / $0.40 | 34.8% | — | — | — | — | — | — | — |
| 91 | GLM 4.7 (partial coverage) | Z.ai | 39.2 | 1 of 8 | $0.40 / $1.75 | — | — | 58.1% | — | — | — | — | — |
| 92 | Gemini 2.5 Flash (partial coverage) | Google | 38.6 | 1 of 8 | $0.30 / $2.50 | 28.7% | — | — | — | — | — | — | — |
| 93 | Qwen3.7 Plus | Qwen | 38.4 | 5 of 8 | $0.32 / $1.28 | — | — | — | -2.0 | -5.6 | -1.1 | 0.2 | 6.3 |
| 94 | Gemini 3.1 Flash Lite (high) (partial coverage) | Google | 38.0 | 1 of 8 | $0.25 / $1.50 | — | — | 57.1% | — | — | — | — | — |
| 95 | gpt-oss-120b (partial coverage) | OpenAI | 37.9 | 1 of 8 | $0.03 / $0.17 | 26.0% | — | — | — | — | — | — | — |
| 96 | MiniMax M3 | MiniMax | 37.5 | 5 of 8 | $0.30 / $1.20 | — | — | — | -2.5 | -5.1 | -5.8 | 0.7 | 5.9 |
| 97 | GPT-4.1 Mini (partial coverage) | OpenAI | 37.1 | 1 of 8 | $0.40 / $1.60 | 23.9% | — | — | — | — | — | — | — |
| 98 | GPT-5.4 Mini (xhigh) (partial coverage) | OpenAI | 36.9 | 1 of 8 | $0.75 / $4.50 | — | — | 56.7% | — | — | — | — | — |
| 99 | GPT-4o (2024-11-20) (partial coverage) | OpenAI | 36.4 | 1 of 8 | $2.50 / $10.00 | 21.6% | — | — | — | — | — | — | — |
| 100 | GPT-5.1 (high) (partial coverage) | OpenAI | 35.7 | 1 of 8 | $1.25 / $10.00 | — | — | 50.1% | — | — | — | — | — |
| 101 | Llama 4 Maverick 17B 128e Instruct (partial coverage) | Meta | 35.7 | 1 of 8 | — | 21.0% | — | — | — | — | — | — | — |
| 102 | DeepSeek V4 Flash 0423 | DeepSeek | 35.5 | 5 of 8 | $0.0643 / $0.1285 | — | — | — | -2.2 | -1.0 | -2.0 | -1.3 | 2.6 |
| 103 | MiMo-V2.5-Pro | Xiaomi | 35.5 | 5 of 8 | $0.435 / $0.87 | — | — | — | -2.0 | -2.0 | -2.8 | 0.0 | 1.6 |
| 104 | Gemini 2.0 Flash 001 (partial coverage) | Google | 34.9 | 1 of 8 | — | 13.5% | — | — | — | — | — | — | — |
| 105 | o3 Pro (partial coverage) | OpenAI | 34.6 | 1 of 8 | $20.00 / $80.00 | — | — | 44.5% | — | — | — | — | — |
| 106 | Llama 4 Scout 17B 16e Instruct (partial coverage) | Meta | 34.2 | 1 of 8 | — | 9.1% | — | — | — | — | — | — | — |
| 107 | Claude Haiku 4.5 (partial coverage) | Anthropic | 33.4 | 1 of 8 | $1.00 / $5.00 | — | — | 40.2% | — | — | — | — | — |
| 108 | Qwen2.5 Coder 32B Instruct (partial coverage) | Qwen | 33.4 | 1 of 8 | — | 9.0% | — | — | — | — | — | — | — |
| 109 | GLM 5.1 (max) (partial coverage) | Z.ai | 33.4 | 1 of 8 | $0.966 / $3.036 | — | 58.7% | — | — | — | — | — | — |
| 110 | Gemini 3.5 Flash (medium) | Google | 33.2 | 5 of 8 | $1.50 / $9.00 | — | — | — | -3.5 | -3.5 | -8.3 | 0.5 | -0.3 |
| 111 | Hy3 | Tencent | 32.5 | 5 of 8 | $0.132 / $0.528 | — | — | — | -1.3 | -8.6 | -2.5 | -1.3 | 4.0 |
| 112 | Inkling | Thinking Machines | 32.2 | 5 of 8 | $0.95 / $4.05 | — | — | — | -6.6 | -11.8 | -11.7 | 0.6 | 6.8 |
| 113 | Grok 4.3 (high) | SpaceXAI | 30.9 | 5 of 8 | $1.25 / $2.50 | — | — | — | -8.5 | -7.2 | -9.7 | 1.1 | -13.3 |
| 114 | Grok Build 0.1 | SpaceXAI | 28.9 | 5 of 8 | $1.00 / $2.00 | — | — | — | -9.1 | -8.6 | -5.4 | 1.0 | -21.4 |
| 115 | Grok 4.3 | SpaceXAI | 27.4 | 5 of 8 | $1.25 / $2.50 | — | — | — | -14.7 | -5.7 | -10.8 | 1.1 | -41.8 |
| 116 | Gemini 3 Flash Preview | Google | 26.2 | 6 of 8 | $0.50 / $3.00 | — | — | 62.0% | -8.5 | -3.4 | -7.2 | -0.7 | -21.0 |
| 117 | Solar Pro 4 | Upstage | 24.9 | 5 of 8 | $0.03 / $0.12 | — | — | — | -10.1 | -13.3 | -8.4 | 0.5 | -13.6 |
| 118 | MiniMax M2.7 | MiniMax | 24.4 | 5 of 8 | $0.30 / $1.20 | — | — | — | -11.2 | -13.7 | -10.9 | 1.0 | -16.8 |
| 119 | Mistral Medium 3.5 | Mistral | 23.8 | 5 of 8 | $1.50 / $7.50 | — | — | — | -7.0 | -11.0 | -9.7 | -2.9 | -1.7 |
| 120 | Gemma 4 31B | Google | 22.7 | 5 of 8 | $0.10 / $0.34 | — | — | — | -18.9 | -9.4 | 0.3 | -31.2 | -51.5 |
| 121 | Gemini 3.5 Flash Lite | Google | 21.5 | 5 of 8 | $0.30 / $2.50 | — | — | — | -10.4 | -10.3 | -12.7 | -0.3 | -15.1 |
| 122 | Nemotron 3 Ultra | NVIDIA | 20.3 | 5 of 8 | $0.60 / $3.60 | — | — | — | -14.6 | -19.9 | -15.4 | 0.5 | -24.1 |

## How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 8 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

## Data sources

- [LMArena](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset): CC BY 4.0. Arena ratings by LMArena, from the public leaderboard dataset.
- [MCP Atlas](https://labs.scale.com/leaderboard/mcp_atlas): MIT. All-1,000-task Pass Rate from Scale's current Performance Comparison, using the standardized MCP loop and 100-tool-call budget.
- [SWE-bench](https://www.swebench.com): Resolve rates published by the SWE-bench maintainers.
- [Terminal-Bench 2.1](https://www.tbench.ai/leaderboard/terminal-bench/2.1): Apache-2.0. Published Accuracy over 89 Terminal-Bench 2.1 tasks, run and verified by the benchmark team; results are model plus agent harness plus effort, not model-only evaluations.
