LLM leaderboard
The overall table is a composite of four text categories: coding, agentic tool use, math and vision. Media categories are reachable by tab and stay out of the overall number, because ranking a speech model against a coding model produces a figure that means nothing.
Coverage is printed on every row. A model the benchmarks have only partly reached still ranks, on the results it does have, with a mark on its row saying so, and its score is pulled toward the mean of the broadly benchmarked models rather than padded with blanks.
Aldena runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | categories | composite | composite | composite | composite |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 68.6 | $5.00 / $25.00 | 3 of 4 categories | 75.9 | 86.1 | 80.5 | |
| 2 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 67.9 | $5.00 / $25.00 | 4 of 4 categories | 78.0 | 82.2 | 65.8 | 81.0 |
| 3 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 65.9 | $10.00 / $50.00 | 4 of 4 categories | 63.3 | 81.8 | 65.9 | 84.0 |
| 4 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 65.8 | $3.00 / $15.00 | 3 of 4 categories | 75.7 | 79.1 | 74.1 | |
| 5 | Claude Fable 5 (max)anthropic/claude-fable-5:max | Anthropic | 65.2 | $10.00 / $50.00 | 2 of 4 categories | 78.7 | 82.0 | ||
| 6 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 65.2 | $5.00 / $25.00 | 4 of 4 categories | 74.0 | 75.8 | 65.3 | 75.9 |
| 7 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 65.2 | $10.00 / $50.00 | 3 of 4 categories | 81.0 | 78.8 | 65.8 | |
| 8 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 65.2 | $5.00 / $30.00 | 4 of 4 categories | 77.1 | 80.4 | 62.5 | 70.7 |
| 9 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 64.3 | $5.00 / $25.00 | 4 of 4 categories | 68.8 | 71.5 | 64.1 | 80.9 |
| 10 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 64.2 | $2.00 / $6.00 | 4 of 4 categories | 69.4 | 74.6 | 62.8 | 78.3 |
| 11 | GPT-5.6 Sol (max)openai/gpt-5.6-sol:max | OpenAI | 64.1 | $5.00 / $30.00 | 2 of 4 categories | 77.0 | 79.3 | ||
| 12 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 64.1 | $2.00 / $6.00 | 4 of 4 categories | 60.2 | 76.7 | 65.5 | 81.7 |
| 13 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 64.0 | $5.00 / $30.00 | 4 of 4 categories | 73.3 | 69.2 | 64.2 | 76.9 |
| 14 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 63.3 | $2.50 / $15.00 | 4 of 4 categories | 64.3 | 66.4 | 78.8 | 70.1 |
| 15 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 63.0 | $5.00 / $30.00 | 3 of 4 categories | 73.3 | 70.5 | 70.9 | |
| 16 | GPT-5.5openai/gpt-5.5 | OpenAI | 62.6 | $5.00 / $30.00 | 4 of 4 categories | 68.8 | 64.3 | 64.8 | 77.6 |
| 17 | Gemini 3 Progoogle/gemini-3-pro | 61.6 | — | 3 of 4 categories | 65.0 | 62.6 | 79.9 | ||
| 18 | Muse Sparkmeta/muse-spark | Meta | 61.4 | — | 4 of 4 categories | 60.5 | 69.1 | 64.4 | 74.4 |
| 19 | GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh | OpenAI | 61.2 | — | 2 of 4 categories | 70.8 | 73.7 | ||
| 20 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 61.1 | $5.00 / $25.00 | 3 of 4 categories | 69.1 | 71.0 | 64.9 | |
| 21 | Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh | Meta | 60.9 | $1.25 / $4.25 | 2 of 4 categories | 68.9 | 74.4 | ||
| 22 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 60.9 | $5.00 / $25.00 | 3 of 4 categories | 60.5 | 68.9 | 74.8 | |
| 23 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 60.8 | $1.25 / $4.25 | 4 of 4 categories | 60.2 | 73.5 | 63.6 | 67.3 |
| 24 | Claude Fable 5 (xhigh)anthropic/claude-fable-5:xhigh | Anthropic | 60.8 | $10.00 / $50.00 | 2 of 4 categories | 66.8 | 76.3 | ||
| 25 | Claude Opus 4.6 (thinking)anthropic/claude-opus-4.6:thinking | Anthropic | 60.5 | $5.00 / $25.00 | 1 of 4 categories | 81.3 | |||
| 26 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 60.4 | $5.00 / $25.00 | 4 of 4 categories | 57.2 | 72.0 | 61.4 | 71.6 |
| 27 | Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking | Anthropic | 60.3 | $5.00 / $25.00 | 1 of 4 categories | 80.5 | |||
| 28 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 60.0 | $1.00 / $6.00 | 4 of 4 categories | 60.2 | 66.2 | 63.4 | 69.8 |
| 29 | Grok 4.6 (high)x-ai/grok-4.6:high | SpaceXAI | 59.9 | $2.00 / $6.00 | 2 of 4 categories | 75.5 | 63.8 | ||
| 30 | Claude Opus 5 (xhigh)anthropic/claude-opus-5:xhigh | Anthropic | 59.7 | $5.00 / $25.00 | 2 of 4 categories | 65.6 | 73.0 | ||
| 31 | GPT-5.6 Terra (max)openai/gpt-5.6-terra:max | OpenAI | 59.5 | $1.00 / $6.00 | 3 of 4 categories | 52.9 | 67.2 | 77.0 | |
| 32 | Gemini 3.7 Flash (high)google/gemini-3.7-flash:high | 59.4 | $0.375 / $1.875 | 3 of 4 categories | 57.8 | 69.8 | 69.2 | ||
| 33 | Claude Opus 4.8 (thinking)anthropic/claude-opus-4.8:thinking | Anthropic | 59.4 | $5.00 / $25.00 | 1 of 4 categories | 77.9 | |||
| 34 | DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high | DeepSeek | 59.0 | $1.168 / $2.336 | 2 of 4 categories | 68.8 | 67.0 | ||
| 35 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 59.0 | $2.00 / $12.00 | 4 of 4 categories | 44.4 | 58.8 | 72.6 | 77.9 | |
| 36 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 58.9 | $2.00 / $10.00 | 4 of 4 categories | 57.1 | 63.5 | 60.4 | 72.0 |
| 37 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 58.6 | $0.50 / $3.00 | 3 of 4 categories | 65.7 | 65.7 | 61.6 | ||
| 38 | Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | 58.6 | $1.027 / $6.162 | 2 of 4 categories | 70.0 | 64.1 | ||
| 39 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 58.6 | $1.75 / $14.00 | 3 of 4 categories | 67.1 | 58.1 | 67.3 | |
| 40 | Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium | Anthropic | 58.5 | $5.00 / $25.00 | 2 of 4 categories | 63.8 | 70.0 | ||
| 41 | Gemini 3.6 Flashgoogle/gemini-3.6-flash | 58.5 | $0.75 / $3.75 | 1 of 4 categories | 75.1 | ||||
| 42 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 58.4 | $1.75 / $14.00 | 4 of 4 categories | 61.6 | 58.3 | 70.6 | 59.7 |
| 43 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 58.3 | $0.462 / $1.452 | 3 of 4 categories | 66.1 | 61.4 | 63.6 | |
| 44 | Claude Opus 5 (medium)anthropic/claude-opus-5:medium | Anthropic | 58.2 | $5.00 / $25.00 | 1 of 4 categories | 74.3 | |||
| 45 | GPT 5.5 Pro Pre Release (xhigh)openai/gpt-5.5-pro-pre-release:xhigh | OpenAI | 58.1 | — | 1 of 4 categories | 74.2 | |||
| 46 | GPT-5.5 Pro (xhigh)openai/gpt-5.5-pro:xhigh | OpenAI | 57.9 | $30.00 / $180.00 | 1 of 4 categories | 73.5 | |||
| 47 | MiniMax M2.5 (high)minimax/minimax-m2.5:high | MiniMax | 57.9 | $0.22 / $0.90 | 2 of 4 categories | 65.7 | 65.7 | ||
| 48 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 57.9 | $0.10 / $0.60 | 4 of 4 categories | 60.8 | 62.8 | 63.5 | 59.9 |
| 49 | GPT-5.4openai/gpt-5.4 | OpenAI | 57.8 | $2.50 / $15.00 | 3 of 4 categories | 57.0 | 59.4 | 72.6 | |
| 50 | Grok 4.6 (xhigh)x-ai/grok-4.6:xhigh | SpaceXAI | 57.8 | $2.00 / $6.00 | 2 of 4 categories | 62.7 | 68.1 | ||
| 51 | Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k | Anthropic | 57.7 | $5.00 / $25.00 | 2 of 4 categories | 70.0 | 60.7 | ||
| 52 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 57.7 | $1.50 / $9.00 | 4 of 4 categories | 47.6 | 56.1 | 72.8 | 69.6 | |
| 53 | GPT-5.4 Pro (xhigh)openai/gpt-5.4-pro:xhigh | OpenAI | 57.6 | $30.00 / $180.00 | 1 of 4 categories | 72.5 | |||
| 54 | Claude Fable 5 (low)anthropic/claude-fable-5:low | Anthropic | 57.5 | $10.00 / $50.00 | 2 of 4 categories | 66.1 | 63.8 | ||
| 55 | Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh | Qwen | 57.4 | $2.00 / $6.00 | 2 of 4 categories | 56.5 | 72.7 | ||
| 56 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 57.3 | $3.00 / $15.00 | 4 of 4 categories | 52.1 | 70.2 | 59.8 | 61.3 |
| 57 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 57.3 | — | 3 of 4 categories | 66.1 | 57.9 | 62.0 | |
| 58 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 57.2 | $2.00 / $6.00 | 3 of 4 categories | 58.4 | 65.7 | 61.8 | |
| 59 | Ernie 5.1baidu/ernie-5.1 | Baidu | 57.2 | — | 2 of 4 categories | 66.7 | 61.9 | ||
| 60 | Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | 57.2 | $5.00 / $25.00 | 2 of 4 categories | 77.3 | 51.4 | ||
| 61 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 57.1 | $5.00 / $25.00 | 3 of 4 categories | 63.1 | 57.6 | 64.4 | |
| 62 | GPT-5.6 Luna (max)openai/gpt-5.6-luna:max | OpenAI | 57.0 | $0.10 / $0.60 | 3 of 4 categories | 47.3 | 63.0 | 74.6 | |
| 63 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 56.9 | $0.0643 / $0.1285 | 3 of 4 categories | 60.3 | 67.6 | 56.4 | |
| 64 | Qwen3.5 Max Previewqwen/qwen3.5-max-preview | Qwen | 56.9 | — | 2 of 4 categories | 66.4 | 60.9 | ||
| 65 | Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high | Anthropic | 56.8 | $5.00 / $25.00 | 2 of 4 categories | 59.6 | 67.4 | ||
| 66 | Claude Fable 5 (medium)anthropic/claude-fable-5:medium | Anthropic | 56.7 | $10.00 / $50.00 | 1 of 4 categories | 69.9 | |||
| 67 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 56.7 | — | 3 of 4 categories | 65.1 | 58.4 | 59.6 | |
| 68 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 56.7 | $0.57 / $2.85 | 3 of 4 categories | 61.5 | 60.8 | 60.8 | |
| 69 | Doubao-Seed-Codebytedance/doubao-seed-code | ByteDance | 56.6 | — | 1 of 4 categories | 69.7 | |||
| 70 | Claude Opus 4.6 (64K)anthropic/claude-opus-4.6:64k | Anthropic | 56.6 | $5.00 / $25.00 | 1 of 4 categories | 69.6 | |||
| 71 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 56.6 | $0.50 / $3.00 | 3 of 4 categories | 55.4 | 58.6 | 68.7 | ||
| 72 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 56.5 | $1.475 / $4.425 | 3 of 4 categories | 45.6 | 66.8 | 69.9 | |
| 73 | GLM-5.3z-ai/glm-5.3 | Z.ai | 56.5 | — | 1 of 4 categories | 69.1 | |||
| 74 | o4 Mini Highopenai/o4-mini-high | OpenAI | 56.5 | $1.10 / $4.40 | 2 of 4 categories | 71.2 | 54.4 | ||
| 75 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 56.3 | — | 3 of 4 categories | 57.6 | 61.3 | 62.4 | ||
| 76 | GPT-5.6 Sol (high)openai/gpt-5.6-sol:high | OpenAI | 56.3 | $5.00 / $30.00 | 1 of 4 categories | 68.6 | |||
| 77 | o3 (high)openai/o3:high | OpenAI | 56.3 | $2.00 / $8.00 | 2 of 4 categories | 69.9 | 54.9 | ||
| 78 | GPT 5.5 Pro Pre Release (high)openai/gpt-5.5-pro-pre-release:high | OpenAI | 56.2 | — | 1 of 4 categories | 68.4 | |||
| 79 | Claude Opus 4.6 (32K)anthropic/claude-opus-4.6:32k | Anthropic | 56.1 | $5.00 / $25.00 | 1 of 4 categories | 67.9 | |||
| 80 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 56.0 | $0.39 / $2.34 | 3 of 4 categories | 58.2 | 61.4 | 60.2 | |
| 81 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 56.0 | — | 3 of 4 categories | 58.3 | 60.1 | 61.4 | |
| 82 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 56.0 | $0.50 / $3.00 | 4 of 4 categories | 26.2 | 72.2 | 64.6 | 72.4 | |
| 83 | Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high | Anthropic | 56.0 | $5.00 / $25.00 | 2 of 4 categories | 57.9 | 65.6 | ||
| 84 | Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k | Anthropic | 55.9 | $15.00 / $75.00 | 2 of 4 categories | 66.0 | 57.4 | ||
| 85 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 55.9 | — | 3 of 4 categories | 66.7 | 44.6 | 67.7 | |
| 86 | Grok 4.6x-ai/grok-4.6 | SpaceXAI | 55.8 | $2.00 / $6.00 | 1 of 4 categories | 67.0 | |||
| 87 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 55.8 | $1.50 / $9.00 | 4 of 4 categories | 33.2 | 66.0 | 62.2 | 72.8 | |
| 88 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 55.7 | $0.5415 / $2.28 | 4 of 4 categories | 47.7 | 52.9 | 67.2 | 66.1 |
| 89 | GLM 5 (high)z-ai/glm-5:high | Z.ai | 55.7 | $0.60 / $1.92 | 2 of 4 categories | 61.6 | 60.8 | ||
| 90 | GPT-5.2 (medium)openai/gpt-5.2:medium | OpenAI | 55.5 | $1.75 / $14.00 | 1 of 4 categories | 66.3 | |||
| 91 | Gemini 3.7 Flash (medium)google/gemini-3.7-flash:medium | 55.5 | $0.375 / $1.875 | 1 of 4 categories | 66.3 | ||||
| 92 | GPT-5.2 Pro (xhigh)openai/gpt-5.2-pro:xhigh | OpenAI | 55.3 | $21.00 / $168.00 | 1 of 4 categories | 65.7 | |||
| 93 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 55.3 | $2.50 / $15.00 | 3 of 4 categories | 48.4 | 55.1 | 72.5 | |
| 94 | Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k | Anthropic | 55.2 | $3.00 / $15.00 | 2 of 4 categories | 61.5 | 58.9 | ||
| 95 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 55.1 | $0.32 / $1.28 | 4 of 4 categories | 38.4 | 64.7 | 65.9 | 61.6 |
| 96 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 55.1 | $5.00 / $25.00 | 3 of 4 categories | 46.5 | 65.5 | 63.3 | |
| 97 | Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | 55.1 | $2.00 / $12.00 | 1 of 4 categories | 65.1 | ||||
| 98 | Kimi K3 (none)moonshotai/kimi-k3:none | MoonshotAI | 55.1 | $3.00 / $15.00 | 1 of 4 categories | 65.0 | |||
| 99 | Claude Opus 5 (low)anthropic/claude-opus-5:low | Anthropic | 55.0 | $5.00 / $25.00 | 2 of 4 categories | 60.2 | 59.5 | ||
| 100 | GPT-5 (high)openai/gpt-5:high | OpenAI | 55.0 | $1.25 / $10.00 | 3 of 4 categories | 56.9 | 65.5 | 52.4 | |
| 101 | Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high | Anthropic | 55.0 | $3.00 / $15.00 | 2 of 4 categories | 60.1 | 59.5 | ||
| 102 | EXAONE 4.0 32Blg-ai/exaone-4.0-32b | LG AI Research | 54.9 | — | 1 of 4 categories | 64.5 | |||
| 103 | Mimo v2 Proxiaomi/mimo-v2-pro | Xiaomi | 54.9 | — | 2 of 4 categories | 61.2 | 58.2 | ||
| 104 | Grok 4.6 (medium)x-ai/grok-4.6:medium | SpaceXAI | 54.9 | $2.00 / $6.00 | 1 of 4 categories | 64.4 | |||
| 105 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 54.9 | $1.25 / $10.00 | 3 of 4 categories | 52.7 | 57.7 | 63.6 | |
| 106 | GPT-5.4 (medium)openai/gpt-5.4:medium | OpenAI | 54.8 | $2.50 / $15.00 | 2 of 4 categories | 57.4 | 61.6 | ||
| 107 | Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview | ByteDance | 54.8 | — | 1 of 4 categories | 64.0 | |||
| 108 | GPT-5.3-Codex (high)openai/gpt-5.3-codex:high | OpenAI | 54.8 | $1.75 / $14.00 | 1 of 4 categories | 64.0 | |||
| 109 | Qwen3 Maxqwen/qwen3-max | Qwen | 54.7 | $0.78 / $3.90 | 2 of 4 categories | 58.4 | 60.1 | ||
| 110 | Claude Opus 5anthropic/claude-opus-5 | Anthropic | 54.7 | $5.00 / $25.00 | 1 of 4 categories | 63.8 | |||
| 111 | GLM 5.1z-ai/glm-5.1 | Z.ai | 54.7 | $0.966 / $3.036 | 3 of 4 categories | 44.0 | 62.9 | 66.2 | |
| 112 | GPT-5openai/gpt-5 | OpenAI | 54.6 | $1.25 / $10.00 | 1 of 4 categories | 63.6 | |||
| 113 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 54.5 | $5.00 / $25.00 | 3 of 4 categories | 53.0 | 50.6 | 68.7 | |
| 114 | DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max | DeepSeek | 54.5 | $1.168 / $2.336 | 2 of 4 categories | 67.5 | 50.2 | ||
| 115 | Kimi K2.5 (high)moonshotai/kimi-k2.5:high | MoonshotAI | 54.5 | $0.57 / $2.85 | 2 of 4 categories | 59.4 | 58.3 | ||
| 116 | OpenReasoning Nemotron 32Bnvidia/openreasoning-nemotron-32b | NVIDIA | 54.5 | — | 1 of 4 categories | 63.2 | |||
| 117 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 54.5 | $1.75 / $14.00 | 3 of 4 categories | 43.8 | 59.4 | 68.8 | |
| 118 | Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max | Anthropic | 54.4 | $2.00 / $10.00 | 2 of 4 categories | 58.7 | 58.6 | ||
| 119 | Ernie 5.0 0110baidu/ernie-5.0-0110 | Baidu | 54.3 | — | 2 of 4 categories | 61.2 | 55.8 | ||
| 120 | GLM 5.2z-ai/glm-5.2 | Z.ai | 54.3 | $0.462 / $1.452 | 2 of 4 categories | 54.1 | 62.9 | ||
| 121 | Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh | Anthropic | 54.2 | $2.00 / $10.00 | 2 of 4 categories | 56.0 | 60.7 | ||
| 122 | Grok 4.1x-ai/grok-4.1 | xAI | 54.2 | — | 2 of 4 categories | 62.1 | 54.4 | ||
| 123 | Claude Opus 4.8 (xhigh)anthropic/claude-opus-4.8:xhigh | Anthropic | 54.1 | $5.00 / $25.00 | 1 of 4 categories | 62.2 | |||
| 124 | GPT-5.6 Luna (high)openai/gpt-5.6-luna:high | OpenAI | 54.1 | $0.10 / $0.60 | 1 of 4 categories | 62.1 | |||
| 125 | GPT-5.3 Chatopenai/gpt-5.3-chat | OpenAI | 54.1 | — | 2 of 4 categories | 62.5 | 53.6 | ||
| 126 | SWE-1.7 (none)cognition/swe-1.7:none | Cognition | 53.9 | — | 1 of 4 categories | 61.5 | |||
| 127 | DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high | DeepSeek | 53.9 | $0.269 / $0.40 | 2 of 4 categories | 57.9 | 57.6 | ||
| 128 | GLM 5z-ai/glm-5 | Z.ai | 53.8 | $0.60 / $1.92 | 2 of 4 categories | 64.1 | 50.9 | ||
| 129 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 53.8 | $1.25 / $10.00 | 4 of 4 categories | 35.7 | 59.9 | 62.3 | 64.7 |
| 130 | Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 53.7 | $1.25 / $10.00 | 1 of 4 categories | 60.9 | ||||
| 131 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 53.7 | $0.12 / $0.40 | 3 of 4 categories | 53.3 | 60.2 | 54.6 | ||
| 132 | Grok 4.5 (medium)x-ai/grok-4.5:medium | SpaceXAI | 53.6 | $2.00 / $6.00 | 1 of 4 categories | 60.7 | |||
| 133 | Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high | 53.6 | — | 2 of 4 categories | 57.1 | 57.0 | |||
| 134 | OpenCodeReasoning Nemotron 1.1 32Bnvidia/opencodereasoning-nemotron-1.1-32b | NVIDIA | 53.6 | — | 1 of 4 categories | 60.5 | |||
| 135 | DeepSeek V4 Flash 0731 (max)deepseek/deepseek-v4-flash-0731:max | DeepSeek | 53.6 | $0.14 / $0.28 | 1 of 4 categories | 60.4 | |||
| 136 | DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking | DeepSeek | 53.5 | $0.27 / $0.41 | 2 of 4 categories | 59.0 | 54.6 | ||
| 137 | Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | 53.4 | $0.325 / $1.95 | 2 of 4 categories | 48.1 | 65.2 | ||
| 138 | o4 Mini (medium)openai/o4-mini:medium | OpenAI | 53.4 | $1.10 / $4.40 | 2 of 4 categories | 68.5 | 44.7 | ||
| 139 | GPT-5 Pro (high)openai/gpt-5-pro:high | OpenAI | 53.3 | $15.00 / $120.00 | 1 of 4 categories | 59.6 | |||
| 140 | Kimi K3 (high)moonshotai/kimi-k3:high | MoonshotAI | 53.3 | $3.00 / $15.00 | 1 of 4 categories | 59.5 | |||
| 141 | GPT-5.6 Sol (medium)openai/gpt-5.6-sol:medium | OpenAI | 53.3 | $5.00 / $30.00 | 1 of 4 categories | 59.5 | |||
| 142 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 53.3 | $0.14 / $0.28 | 3 of 4 categories | 60.6 | 56.6 | 48.8 | |
| 143 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 53.2 | $0.75 / $3.75 | 3 of 4 categories | 40.2 | 59.1 | 66.6 | ||
| 144 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 53.1 | — | 3 of 4 categories | 60.4 | 56.1 | 49.0 | |
| 145 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 53.1 | $0.435 / $0.87 | 3 of 4 categories | 35.5 | 67.6 | 62.1 | |
| 146 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 53.0 | $0.75 / $4.50 | 3 of 4 categories | 50.1 | 57.6 | 57.1 | |
| 147 | GPT-5.6 Solopenai/gpt-5.6-sol | OpenAI | 53.0 | $5.00 / $30.00 | 1 of 4 categories | 58.7 | |||
| 148 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Meta | 52.9 | $0.35 / $1.50 | 2 of 4 categories | 52.9 | 58.4 | ||
| 149 | Hy3tencent/hy3 | Tencent | 52.8 | $0.132 / $0.528 | 3 of 4 categories | 32.5 | 67.8 | 63.2 | |
| 150 | Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high | Anthropic | 52.7 | $1.00 / $5.00 | 2 of 4 categories | 54.9 | 55.7 | ||
| 151 | Qwen3.6 27Bqwen/qwen3.6-27b | Qwen | 52.7 | $0.60 / $3.60 | 1 of 4 categories | 57.9 | |||
| 152 | Kimi K2 0905moonshotai/kimi-k2-0905 | MoonshotAI | 52.6 | $0.60 / $2.50 | 2 of 4 categories | 58.7 | 51.5 | ||
| 153 | Inkling Smallthinkingmachines/inkling-small | Thinking Machines | 52.6 | $0.45 / $1.20 | 1 of 4 categories | 57.6 | |||
| 154 | Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | Qwen | 52.6 | $0.09 / $0.55 | 2 of 4 categories | 58.3 | 51.8 | ||
| 155 | GPT-5.2openai/gpt-5.2 | OpenAI | 52.6 | $1.75 / $14.00 | 4 of 4 categories | 56.4 | 57.4 | 54.9 | 46.5 |
| 156 | R1 0528deepseek/deepseek-r1-0528 | DeepSeek | 52.6 | $0.50 / $2.15 | 2 of 4 categories | 63.7 | 46.3 | ||
| 157 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 52.5 | — | 3 of 4 categories | 60.8 | 55.2 | 46.5 | |
| 158 | LongCat Flash Chatmeituan/longcat-flash-chat | Meituan | 52.5 | — | 2 of 4 categories | 58.7 | 50.9 | ||
| 159 | o3openai/o3 | OpenAI | 52.4 | $2.00 / $8.00 | 4 of 4 categories | 48.3 | 61.5 | 57.6 | 46.8 |
| 160 | Gemini 3.5 Flash (low)google/gemini-3.5-flash:low | 52.4 | $1.50 / $9.00 | 1 of 4 categories | 56.8 | ||||
| 161 | gpt-oss-120b (high)openai/gpt-oss-120b:high | OpenAI | 52.4 | $0.03 / $0.17 | 1 of 4 categories | 56.8 | |||
| 162 | DeepSeek V3.1 Terminus (thinking)deepseek/deepseek-v3.1-terminus:thinking | DeepSeek | 52.3 | $0.27 / $0.95 | 1 of 4 categories | 56.6 | |||
| 163 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 52.3 | $1.20 / $4.00 | 3 of 4 categories | 57.6 | 57.2 | 46.4 | |
| 164 | GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium | OpenAI | 52.3 | $1.25 / $10.00 | 2 of 4 categories | 53.8 | 55.1 | ||
| 165 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 52.2 | — | 3 of 4 categories | 57.8 | 48.6 | 54.5 | |
| 166 | GPT-5.6 Sol (low)openai/gpt-5.6-sol:low | OpenAI | 52.2 | $5.00 / $30.00 | 2 of 4 categories | 47.0 | 61.6 | ||
| 167 | Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh | Anthropic | 52.1 | $5.00 / $25.00 | 2 of 4 categories | 51.9 | 56.3 | ||
| 168 | Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low | Anthropic | 52.1 | $5.00 / $25.00 | 2 of 4 categories | 44.3 | 63.8 | ||
| 169 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 52.1 | $0.25 / $1.50 | 3 of 4 categories | 44.5 | 55.6 | 60.0 | ||
| 170 | Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203 | Baidu | 52.1 | — | 2 of 4 categories | 58.2 | 49.9 | ||
| 171 | GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | 52.0 | $1.75 / $14.00 | 2 of 4 categories | 61.6 | 46.3 | ||
| 172 | Grok 4.5 (low)x-ai/grok-4.5:low | SpaceXAI | 52.0 | $2.00 / $6.00 | 1 of 4 categories | 55.8 | |||
| 173 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 52.0 | $0.25 / $2.00 | 3 of 4 categories | 49.0 | 54.1 | 56.8 | |
| 174 | o1 (medium)openai/o1:medium | OpenAI | 52.0 | $15.00 / $60.00 | 1 of 4 categories | 55.7 | |||
| 175 | GPT-5.2 (low)openai/gpt-5.2:low | OpenAI | 51.9 | $1.75 / $14.00 | 1 of 4 categories | 55.5 | |||
| 176 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 51.9 | $1.25 / $10.00 | 3 of 4 categories | 53.8 | 51.9 | 53.7 | |
| 177 | Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | Qwen | 51.8 | $0.15 / $1.00 | 1 of 4 categories | 55.3 | |||
| 178 | Qwen3.7 Flashqwen/qwen3.7-flash | Qwen | 51.8 | $0.03 / $0.13 | 1 of 4 categories | 55.3 | |||
| 179 | XBai o4 (medium)metastone/xbai-o4:medium | MetaStoneTec | 51.8 | — | 1 of 4 categories | 55.2 | |||
| 180 | Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | Moonshot AI | 51.8 | — | 2 of 4 categories | 50.4 | 56.5 | ||
| 181 | Qwen3.5 Plusqwen/qwen3.5-plus | Qwen | 51.7 | $0.30 / $1.80 | 1 of 4 categories | 54.8 | |||
| 182 | MiniMax M3minimax/minimax-m3 | MiniMax | 51.7 | $0.30 / $1.20 | 4 of 4 categories | 37.5 | 64.8 | 53.6 | 53.7 |
| 183 | Claude Opus 4.5 (128K)anthropic/claude-opus-4.5:128k | Anthropic | 51.6 | $5.00 / $25.00 | 1 of 4 categories | 54.5 | |||
| 184 | Claude Sonnet 4.6 (32K)anthropic/claude-sonnet-4.6:32k | Anthropic | 51.5 | $3.00 / $15.00 | 1 of 4 categories | 54.4 | |||
| 185 | GPT-5.3-Codexopenai/gpt-5.3-codex | OpenAI | 51.5 | $1.75 / $14.00 | 1 of 4 categories | 54.4 | |||
| 186 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 51.5 | — | 3 of 4 categories | 56.5 | 50.3 | 50.6 | |
| 187 | DeepSeek V4 Pro (xhigh)deepseek/deepseek-v4-pro:xhigh | DeepSeek | 51.5 | $1.168 / $2.336 | 1 of 4 categories | 54.4 | |||
| 188 | Grok 4.1 (thinking)x-ai/grok-4.1:thinking | xAI | 51.5 | — | 2 of 4 categories | 48.9 | 56.9 | ||
| 189 | DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking | DeepSeek | 51.5 | $0.25 / $0.95 | 2 of 4 categories | 55.0 | 50.8 | ||
| 190 | o3 (medium)openai/o3:medium | OpenAI | 51.4 | $2.00 / $8.00 | 2 of 4 categories | 52.6 | 52.6 | ||
| 191 | Claude Opus 4.8 (none)anthropic/claude-opus-4.8:none | Anthropic | 51.3 | $5.00 / $25.00 | 1 of 4 categories | 53.5 | |||
| 192 | GPT-5.4 (low)openai/gpt-5.4:low | OpenAI | 51.3 | $2.50 / $15.00 | 1 of 4 categories | 53.5 | |||
| 193 | Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Qwen | 51.2 | $0.10 / $1.10 | 2 of 4 categories | 53.4 | 51.3 | ||
| 194 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 51.2 | $0.71 / $3.50 | 3 of 4 categories | 52.7 | 47.8 | 55.3 | |
| 195 | GPT-5.5 (medium)openai/gpt-5.5:medium | OpenAI | 51.1 | $5.00 / $30.00 | 1 of 4 categories | 53.2 | |||
| 196 | DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepSeek | 51.1 | $0.25 / $0.95 | 2 of 4 categories | 53.6 | 50.6 | ||
| 197 | R1deepseek/deepseek-r1 | DeepSeek | 51.1 | $0.70 / $2.50 | 2 of 4 categories | 53.1 | 51.1 | ||
| 198 | Inkling Small (xhigh)thinkingmachines/inkling-small:xhigh | Thinking Machines | 51.1 | $0.45 / $1.20 | 1 of 4 categories | 53.0 | |||
| 199 | Step 3.5 Flashstepfun/step-3.5-flash | StepFun | 50.9 | $0.10 / $0.30 | 2 of 4 categories | 54.2 | 49.4 | ||
| 200 | Gemini 3.7 Flash (low)google/gemini-3.7-flash:low | 50.9 | $0.375 / $1.875 | 1 of 4 categories | 52.6 | ||||
| 201 | Grok 4.20 0309 (reasoning)x-ai/grok-4.20-0309:reasoning | xAI | 50.9 | — | 1 of 4 categories | 52.6 | |||
| 202 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 50.9 | $0.26 / $1.04 | 3 of 4 categories | 57.1 | 49.6 | 47.7 | |
| 203 | o1 (high)openai/o1:high | OpenAI | 50.9 | $15.00 / $60.00 | 1 of 4 categories | 52.4 | |||
| 204 | Gemini 2.5 Progoogle/gemini-2.5-pro | 50.9 | $1.25 / $10.00 | 4 of 4 categories | 43.1 | 50.6 | 47.7 | 63.8 | |
| 205 | Qwen3.5 397B A17B (none)qwen/qwen3.5-397b-a17b:none | Qwen | 50.9 | $0.39 / $2.34 | 1 of 4 categories | 52.3 | |||
| 206 | GPT-5.6 Terra (high)openai/gpt-5.6-terra:high | OpenAI | 50.8 | $1.00 / $6.00 | 1 of 4 categories | 52.2 | |||
| 207 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 50.8 | $15.00 / $75.00 | 3 of 4 categories | 63.5 | 52.2 | 38.1 | |
| 208 | o3 Mini (medium)openai/o3-mini:medium | OpenAI | 50.8 | $1.10 / $4.40 | 1 of 4 categories | 52.1 | |||
| 209 | GPT-5.5 (low)openai/gpt-5.5:low | OpenAI | 50.8 | $5.00 / $30.00 | 2 of 4 categories | 49.3 | 53.5 | ||
| 210 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 50.7 | $3.00 / $15.00 | 3 of 4 categories | 58.6 | 54.5 | 40.3 | |
| 211 | Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview | Tencent | 50.7 | — | 2 of 4 categories | 49.4 | 53.2 | ||
| 212 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 50.7 | $1.168 / $2.336 | 3 of 4 categories | 41.7 | 53.9 | 57.5 | |
| 213 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 50.6 | $0.269 / $0.40 | 3 of 4 categories | 49.7 | 50.1 | 53.1 | |
| 214 | Gemini 3.6 Flash (medium)google/gemini-3.6-flash:medium | 50.6 | $0.75 / $3.75 | 1 of 4 categories | 51.4 | ||||
| 215 | Claude Opus 4.5 (16K)anthropic/claude-opus-4.5:16k | Anthropic | 50.6 | $5.00 / $25.00 | 1 of 4 categories | 51.4 | |||
| 216 | DeepSeek V4 Flash 0423 (max)deepseek/deepseek-v4-flash:max | DeepSeek | 50.5 | $0.0643 / $0.1285 | 1 of 4 categories | 51.1 | |||
| 217 | Gemini 3.5 Flash (minimal)google/gemini-3.5-flash:minimal | 50.5 | $1.50 / $9.00 | 1 of 4 categories | 51.1 | ||||
| 218 | Gemini 3.6 Flash (minimal)google/gemini-3.6-flash:minimal | 50.5 | $0.75 / $3.75 | 1 of 4 categories | 51.1 | ||||
| 219 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 50.4 | — | 3 of 4 categories | 55.5 | 43.9 | 52.5 | |
| 220 | Muse Spark 1.1 (xhigh)meta/muse-spark-1.1:xhigh | Meta | 50.4 | $1.25 / $4.25 | 2 of 4 categories | 50.1 | 51.1 | ||
| 221 | o4 Mini (low)openai/o4-mini:low | OpenAI | 50.3 | $1.10 / $4.40 | 2 of 4 categories | 57.2 | 43.9 | ||
| 222 | Claude Sonnet 4.6 (16K)anthropic/claude-sonnet-4.6:16k | Anthropic | 50.3 | $3.00 / $15.00 | 1 of 4 categories | 50.6 | |||
| 223 | GPT-5.1openai/gpt-5.1 | OpenAI | 50.3 | $1.25 / $10.00 | 3 of 4 categories | 50.7 | 52.5 | 47.9 | |
| 224 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 50.2 | $0.30 / $2.50 | 4 of 4 categories | 21.5 | 63.4 | 52.9 | 63.2 | |
| 225 | Kimi K2p5moonshotai/kimi-k2p5 | Moonshot AI | 50.1 | — | 2 of 4 categories | 42.6 | 57.6 | ||
| 226 | Kimi K2.5moonshotai/kimi-k2.5 | MoonshotAI | 50.1 | $0.57 / $2.85 | 1 of 4 categories | 50.1 | |||
| 227 | Kimi K2 0711moonshotai/kimi-k2 | MoonshotAI | 50.1 | $0.57 / $2.30 | 2 of 4 categories | 54.6 | 45.6 | ||
| 228 | Qwen3.7 Flash (none)qwen/qwen3.7-flash:none | Qwen | 50.1 | $0.03 / $0.13 | 1 of 4 categories | 50.0 | |||
| 229 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 50.1 | $0.29 / $2.40 | 3 of 4 categories | 49.4 | 52.6 | 48.1 | |
| 230 | Claude Opus 4 (thinking)anthropic/claude-opus-4:thinking | Anthropic | 50.0 | $15.00 / $75.00 | 1 of 4 categories | 49.9 | |||
| 231 | Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking | Qwen | 50.0 | $0.455 / $1.82 | 2 of 4 categories | 53.2 | 46.5 | ||
| 232 | Kimi K2.7 Code (none)moonshotai/kimi-k2.7-code:none | MoonshotAI | 50.0 | $0.71 / $3.50 | 1 of 4 categories | 49.7 | |||
| 233 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 50.0 | $3.00 / $15.00 | 3 of 4 categories | 58.6 | 48.1 | 43.0 | |
| 234 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 49.9 | $0.0643 / $0.1285 | 3 of 4 categories | 35.5 | 60.4 | 53.3 | |
| 235 | o3 Mini Highopenai/o3-mini-high | OpenAI | 49.9 | $1.10 / $4.40 | 2 of 4 categories | 56.4 | 42.8 | ||
| 236 | Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b | NVIDIA | 49.8 | — | 1 of 4 categories | 49.2 | |||
| 237 | Claude Sonnet 4.5 (59K)anthropic/claude-sonnet-4.5:59k | Anthropic | 49.8 | $3.00 / $15.00 | 1 of 4 categories | 49.1 | |||
| 238 | DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | DeepSeek | 49.8 | $0.27 / $0.95 | 2 of 4 categories | 52.1 | 46.6 | ||
| 239 | Gemma 4 31B (minimal)google/gemma-4-31b-it:minimal | 49.7 | $0.10 / $0.34 | 1 of 4 categories | 48.9 | ||||
| 240 | Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Qwen | 49.6 | $0.23 / $2.30 | 2 of 4 categories | 52.5 | 45.8 | ||
| 241 | Claude Opus 4.1anthropic/claude-opus-4.1 | Anthropic | 49.6 | $15.00 / $75.00 | 2 of 4 categories | 53.9 | 44.3 | ||
| 242 | Inkling (xhigh)thinkingmachines/inkling:xhigh | Thinking Machines | 49.6 | $0.95 / $4.05 | 2 of 4 categories | 51.8 | 46.4 | ||
| 243 | Claude Sonnet 5anthropic/claude-sonnet-5 | Anthropic | 49.6 | $2.00 / $10.00 | 1 of 4 categories | 48.6 | |||
| 244 | Claude Sonnet 4 (thinking)anthropic/claude-sonnet-4:thinking | Anthropic | 49.6 | $3.00 / $15.00 | 1 of 4 categories | 48.5 | |||
| 245 | DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | DeepSeek | 49.5 | $0.27 / $0.41 | 2 of 4 categories | 46.7 | 51.2 | ||
| 246 | Claude Opus 4.7 (low)anthropic/claude-opus-4.7:low | Anthropic | 49.5 | $5.00 / $25.00 | 1 of 4 categories | 48.4 | |||
| 247 | MiniMax M2.5minimax/minimax-m2.5 | MiniMax | 49.5 | $0.22 / $0.90 | 2 of 4 categories | 51.0 | 46.9 | ||
| 248 | Composer 2.5cursor/composer-2.5 | Cursor | 49.5 | — | 1 of 4 categories | 48.3 | |||
| 249 | Claude Sonnet 4 (32K)anthropic/claude-sonnet-4:32k | Anthropic | 49.5 | $3.00 / $15.00 | 1 of 4 categories | 48.1 | |||
| 250 | Claude Sonnet 4.5 (16K)anthropic/claude-sonnet-4.5:16k | Anthropic | 49.5 | $3.00 / $15.00 | 1 of 4 categories | 48.1 | |||
| 251 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 49.4 | $0.60 / $2.50 | 3 of 4 categories | 51.2 | 53.4 | 42.3 | |
| 252 | Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022 | Baidu | 49.4 | — | 2 of 4 categories | 50.3 | 47.1 | ||
| 253 | Claude Opus 4.5 (32K)anthropic/claude-opus-4.5:32k | Anthropic | 49.4 | $5.00 / $25.00 | 1 of 4 categories | 47.9 | |||
| 254 | Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 49.4 | $1.25 / $10.00 | 1 of 4 categories | 47.9 | ||||
| 255 | Devstral Small 2512mistralai/devstral-small-2512 | Mistral AI | 49.3 | — | 2 of 4 categories | 47.5 | 49.6 | ||
| 256 | Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | Qwen | 49.3 | $0.07 / $0.28 | 1 of 4 categories | 47.7 | |||
| 257 | Claude Sonnet 4.5 (32K)anthropic/claude-sonnet-4.5:32k | Anthropic | 49.3 | $3.00 / $15.00 | 1 of 4 categories | 47.5 | |||
| 258 | Gemma 4 31Bgoogle/gemma-4-31b-it | 49.2 | $0.10 / $0.34 | 4 of 4 categories | 22.7 | 56.0 | 61.2 | 55.3 | |
| 259 | Claude Haiku 4.5 (32K)anthropic/claude-haiku-4.5:32k | Anthropic | 49.2 | $1.00 / $5.00 | 1 of 4 categories | 47.4 | |||
| 260 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 49.2 | $0.0482 / $0.1931 | 2 of 4 categories | 52.3 | 44.3 | ||
| 261 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 49.2 | $0.25 / $2.00 | 3 of 4 categories | 50.1 | 57.9 | 37.6 | |
| 262 | Kimi K3 (low)moonshotai/kimi-k3:low | MoonshotAI | 49.1 | $3.00 / $15.00 | 1 of 4 categories | 47.0 | |||
| 263 | GPT-5.4 Nano (low)openai/gpt-5.4-nano:low | OpenAI | 49.1 | $0.20 / $1.25 | 1 of 4 categories | 47.0 | |||
| 264 | GPT-5.6 Sol (none)openai/gpt-5.6-sol:none | OpenAI | 49.1 | $5.00 / $30.00 | 1 of 4 categories | 47.0 | |||
| 265 | Qwen3.6 35B A3B (none)qwen/qwen3.6-35b-a3b:none | Qwen | 49.1 | $0.15 / $1.00 | 1 of 4 categories | 47.0 | |||
| 266 | Devstral Small 2505mistralai/devstral-small-2505 | Mistral AI | 49.1 | — | 1 of 4 categories | 46.9 | |||
| 267 | Laguna M.1poolside/laguna-m.1 | Poolside | 49.0 | — | 1 of 4 categories | 46.9 | |||
| 268 | Gemini 3.6 Flash (low)google/gemini-3.6-flash:low | 49.0 | $0.75 / $3.75 | 2 of 4 categories | 43.6 | 52.3 | |||
| 269 | GPT-5.1 (low)openai/gpt-5.1:low | OpenAI | 49.0 | $1.25 / $10.00 | 1 of 4 categories | 46.7 | |||
| 270 | Composer 2.5 (none)cursor/composer-2.5:none | Cursor | 49.0 | — | 1 of 4 categories | 46.6 | |||
| 271 | GLM 4.5 Airz-ai/glm-4.5-air | Z.ai | 49.0 | $0.13 / $0.85 | 2 of 4 categories | 49.6 | 45.9 | ||
| 272 | Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium | Anthropic | 48.9 | $3.00 / $15.00 | 2 of 4 categories | 43.1 | 52.3 | ||
| 273 | Qwen3 32Bqwen/qwen3-32b | Qwen | 48.9 | $0.08 / $0.28 | 2 of 4 categories | 47.7 | 47.6 | ||
| 274 | Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Qwen | 48.9 | $0.15 / $1.20 | 2 of 4 categories | 49.1 | 46.2 | ||
| 275 | MiniMax M2.1minimax/minimax-m2.1 | MiniMax | 48.9 | $0.30 / $1.20 | 2 of 4 categories | 49.2 | 46.1 | ||
| 276 | Hunyuan T1tencent/hunyuan-t1 | Tencent | 48.8 | — | 2 of 4 categories | 47.2 | 47.9 | ||
| 277 | Qwen3.6 27B (none)qwen/qwen3.6-27b:none | Qwen | 48.8 | $0.60 / $3.60 | 1 of 4 categories | 46.2 | |||
| 278 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 48.8 | $0.20 / $1.25 | 3 of 4 categories | 56.0 | 52.9 | 34.7 | |
| 279 | Inklingthinkingmachines/inkling | Thinking Machines | 48.7 | $0.95 / $4.05 | 3 of 4 categories | 32.2 | 48.0 | 62.9 | |
| 280 | Trinity Large Previewarcee-ai/trinity-large-preview | Arcee AI | 48.7 | — | 2 of 4 categories | 52.7 | 41.8 | ||
| 281 | Grok 3 Mini Beta (low)x-ai/grok-3-mini-beta:low | xAI | 48.7 | — | 1 of 4 categories | 45.7 | |||
| 282 | GPT-5.1-Codexopenai/gpt-5.1-codex | OpenAI | 48.6 | $1.25 / $10.00 | 1 of 4 categories | 45.7 | |||
| 283 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 48.6 | $0.195 / $1.56 | 3 of 4 categories | 47.9 | 54.2 | 40.8 | |
| 284 | Nova Premier 1.0amazon/nova-premier-v1 | Amazon | 48.6 | $2.50 / $12.50 | 1 of 4 categories | 45.6 | |||
| 285 | Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max | Anthropic | 48.5 | $3.00 / $15.00 | 2 of 4 categories | 45.8 | 48.1 | ||
| 286 | Qwen3.6 Flashqwen/qwen3.6-flash | Qwen | 48.5 | $0.1875 / $1.125 | 1 of 4 categories | 45.2 | |||
| 287 | Grok 4.6 (low)x-ai/grok-4.6:low | SpaceXAI | 48.5 | $2.00 / $6.00 | 1 of 4 categories | 45.2 | |||
| 288 | GPT-5.2 (none)openai/gpt-5.2:none | OpenAI | 48.4 | $1.75 / $14.00 | 1 of 4 categories | 45.0 | |||
| 289 | Claude Opus 4.6 (medium)anthropic/claude-opus-4.6:medium | Anthropic | 48.4 | $5.00 / $25.00 | 1 of 4 categories | 44.9 | |||
| 290 | Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high | xAI | 48.4 | — | 2 of 4 categories | 50.6 | 42.6 | ||
| 291 | Claude 3.7 Sonnet (32K)anthropic/claude-3-7-sonnet:32k | Anthropic | 48.3 | — | 1 of 4 categories | 44.7 | |||
| 292 | INTELLECT-3prime-intellect/intellect-3 | Prime Intellect | 48.3 | — | 2 of 4 categories | 48.1 | 44.9 | ||
| 293 | Claude Opus 4 (16K)anthropic/claude-opus-4:16k | Anthropic | 48.3 | $15.00 / $75.00 | 1 of 4 categories | 44.6 | |||
| 294 | Gemini 3.5 Flash Lite (low)google/gemini-3.5-flash-lite:low | 48.3 | $0.30 / $2.50 | 1 of 4 categories | 44.6 | ||||
| 295 | Claude Sonnet 5 (medium)anthropic/claude-sonnet-5:medium | Anthropic | 48.3 | $2.00 / $10.00 | 1 of 4 categories | 44.6 | |||
| 296 | o1openai/o1 | OpenAI | 48.3 | $15.00 / $60.00 | 3 of 4 categories | 50.0 | 43.9 | 47.2 | |
| 297 | Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | NVIDIA | 48.3 | — | 2 of 4 categories | 47.5 | 45.3 | ||
| 298 | GLM 4.7 Flashz-ai/glm-4.7-flash | Z.ai | 48.2 | $0.06 / $0.40 | 2 of 4 categories | 49.5 | 42.9 | ||
| 299 | Laguna XS.2poolside/laguna-xs.2 | Poolside | 48.1 | — | 1 of 4 categories | 44.2 | |||
| 300 | GPT-5.4 (none)openai/gpt-5.4:none | OpenAI | 48.1 | $2.50 / $15.00 | 1 of 4 categories | 44.0 | |||
| 301 | GPT-5.5 (none)openai/gpt-5.5:none | OpenAI | 48.1 | $5.00 / $30.00 | 1 of 4 categories | 44.0 | |||
| 302 | GPT-5.6 Terra (low)openai/gpt-5.6-terra:low | OpenAI | 48.1 | $1.00 / $6.00 | 2 of 4 categories | 35.3 | 56.8 | ||
| 303 | MiniMax M1minimax/minimax-m1 | MiniMax | 48.1 | $0.55 / $2.20 | 2 of 4 categories | 48.7 | 43.3 | ||
| 304 | Devstral Small 2507mistralai/devstral-small-2507 | Mistral AI | 48.1 | — | 1 of 4 categories | 43.9 | |||
| 305 | Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | NVIDIA | 48.0 | $0.085 / $0.40 | 2 of 4 categories | 47.9 | 44.0 | ||
| 306 | GLM 4.7z-ai/glm-4.7 | Z.ai | 48.0 | $0.40 / $1.75 | 3 of 4 categories | 39.2 | 58.7 | 42.1 | |
| 307 | DeepSeek V3.2deepseek/deepseek-v3.2 | DeepSeek | 48.0 | $0.269 / $0.40 | 2 of 4 categories | 40.6 | 51.2 | ||
| 308 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 48.0 | $0.40 / $4.00 | 3 of 4 categories | 54.6 | 48.5 | 36.7 | |
| 309 | DeepSeek V4 Flash 0423 (xhigh)deepseek/deepseek-v4-flash:xhigh | DeepSeek | 48.0 | $0.0643 / $0.1285 | 1 of 4 categories | 43.8 | |||
| 310 | o3 Mini (low)openai/o3-mini:low | OpenAI | 48.0 | $1.10 / $4.40 | 2 of 4 categories | 51.2 | 40.6 | ||
| 311 | Mercuryinception/mercury | Inception | 48.0 | — | 1 of 4 categories | 43.8 | |||
| 312 | Claude Opus 4.1 (27K)anthropic/claude-opus-4.1:27k | Anthropic | 48.0 | $15.00 / $75.00 | 1 of 4 categories | 43.7 | |||
| 313 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 48.0 | $1.25 / $2.50 | 4 of 4 categories | 27.4 | 52.7 | 51.9 | 55.6 |
| 314 | GPT-5 Mini (minimal)openai/gpt-5-mini:minimal | OpenAI | 47.9 | $0.25 / $2.00 | 1 of 4 categories | 43.6 | |||
| 315 | Ernie 5.0 Preview 1220baidu/ernie-5.0-preview-1220 | Baidu | 47.9 | — | 1 of 4 categories | 43.4 | |||
| 316 | o3 (low)openai/o3:low | OpenAI | 47.9 | $2.00 / $8.00 | 1 of 4 categories | 43.4 | |||
| 317 | Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | NVIDIA | 47.9 | — | 1 of 4 categories | 43.3 | |||
| 318 | gpt-oss-20b (high)openai/gpt-oss-20b:high | OpenAI | 47.9 | $0.03 / $0.13 | 1 of 4 categories | 43.3 | |||
| 319 | Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1 | NVIDIA | 47.8 | — | 2 of 4 categories | 46.5 | 44.5 | ||
| 320 | Claude 3.7 Sonnet (16K)anthropic/claude-3-7-sonnet:16k | Anthropic | 47.8 | — | 1 of 4 categories | 43.0 | |||
| 321 | Grok 3 Betax-ai/grok-3-beta | xAI | 47.8 | — | 2 of 4 categories | 52.8 | 37.9 | ||
| 322 | o3 Miniopenai/o3-mini | OpenAI | 47.7 | $1.10 / $4.40 | 2 of 4 categories | 45.8 | 44.8 | ||
| 323 | Claude Sonnet 4 (16K)anthropic/claude-sonnet-4:16k | Anthropic | 47.7 | $3.00 / $15.00 | 1 of 4 categories | 42.9 | |||
| 324 | GPT-5.6 Terra (none)openai/gpt-5.6-terra:none | OpenAI | 47.7 | $1.00 / $6.00 | 1 of 4 categories | 42.9 | |||
| 325 | o1 (low)openai/o1:low | OpenAI | 47.7 | $15.00 / $60.00 | 1 of 4 categories | 42.9 | |||
| 326 | Qwen3.7 Plus (none)qwen/qwen3.7-plus:none | Qwen | 47.6 | $0.32 / $1.28 | 2 of 4 categories | 39.2 | 51.1 | ||
| 327 | GPT-5.4 Mini (medium)openai/gpt-5.4-mini:medium | OpenAI | 47.6 | $0.75 / $4.50 | 1 of 4 categories | 42.7 | |||
| 328 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 47.5 | — | 3 of 4 categories | 45.7 | 47.9 | 43.9 | |
| 329 | Mimo v2 Flashxiaomi/mimo-v2-flash | Xiaomi | 47.5 | — | 2 of 4 categories | 45.8 | 44.2 | ||
| 330 | Gemini 3.5 Flash Lite (minimal)google/gemini-3.5-flash-lite:minimal | 47.5 | $0.30 / $2.50 | 1 of 4 categories | 42.4 | ||||
| 331 | o4 Miniopenai/o4-mini | OpenAI | 47.5 | $1.10 / $4.40 | 4 of 4 categories | 41.6 | 49.8 | 51.1 | 42.5 |
| 332 | KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1 | Kwaipilot | 47.5 | — | 1 of 4 categories | 42.4 | |||
| 333 | DeepSeek V4 Flash 0731 (high)deepseek/deepseek-v4-flash-0731:high | DeepSeek | 47.5 | $0.14 / $0.28 | 1 of 4 categories | 42.2 | |||
| 334 | O1 Mini (high)openai/o1-mini:high | OpenAI | 47.5 | — | 1 of 4 categories | 42.2 | |||
| 335 | Qwen 2.5 (max)qwen/qwen-2.5:max | Qwen | 47.5 | — | 2 of 4 categories | 47.3 | 42.3 | ||
| 336 | Qwen VL (max)qwen/qwen-vl:max | Qwen | 47.4 | — | 1 of 4 categories | 42.0 | |||
| 337 | Qwen3 235B A22Bqwen/qwen3-235b-a22b | Qwen | 47.4 | $0.455 / $1.82 | 2 of 4 categories | 52.5 | 36.8 | ||
| 338 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Qwen | 47.4 | $0.225 / $1.80 | 2 of 4 categories | 41.5 | 47.8 | ||
| 339 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 47.4 | — | 3 of 4 categories | 54.3 | 45.2 | 37.1 | |
| 340 | Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high | 47.4 | $0.25 / $1.50 | 2 of 4 categories | 38.0 | 51.1 | |||
| 341 | Ling Flash 2.0inclusionai/ling-flash-2.0 | inclusionAI | 47.4 | — | 2 of 4 categories | 48.3 | 40.9 | ||
| 342 | Claude Opus 4.6 (low)anthropic/claude-opus-4.6:low | Anthropic | 47.4 | $5.00 / $25.00 | 1 of 4 categories | 41.8 | |||
| 343 | Gemini 3 Pro (high)google/gemini-3-pro:high | 47.3 | — | 1 of 4 categories | 41.8 | ||||
| 344 | Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning | xAI | 47.3 | — | 2 of 4 categories | 39.0 | 50.1 | ||
| 345 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 47.3 | $0.60 / $3.60 | 3 of 4 categories | 20.3 | 59.3 | 56.8 | |
| 346 | GPT-5.6 Terra (medium)openai/gpt-5.6-terra:medium | OpenAI | 47.3 | $1.00 / $6.00 | 1 of 4 categories | 41.6 | |||
| 347 | GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh | OpenAI | 47.3 | $0.75 / $4.50 | 3 of 4 categories | 36.9 | 47.9 | 51.4 | |
| 348 | GPT-5 Nano (low)openai/gpt-5-nano:low | OpenAI | 47.2 | $0.05 / $0.40 | 1 of 4 categories | 41.4 | |||
| 349 | GPT-5 (minimal)openai/gpt-5:minimal | OpenAI | 47.2 | $1.25 / $10.00 | 1 of 4 categories | 41.4 | |||
| 350 | GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | OpenAI | 47.2 | $0.25 / $2.00 | 1 of 4 categories | 41.5 | |||
| 351 | GPT-5.4 Nano (none)openai/gpt-5.4-nano:none | OpenAI | 47.2 | $0.20 / $1.25 | 1 of 4 categories | 41.4 | |||
| 352 | Grok 3 Mini Betax-ai/grok-3-mini-beta | xAI | 47.2 | — | 2 of 4 categories | 45.4 | 43.0 | ||
| 353 | Claude Opus 4 (27K)anthropic/claude-opus-4:27k | Anthropic | 47.2 | $15.00 / $75.00 | 1 of 4 categories | 41.3 | |||
| 354 | Grok 4 0709x-ai/grok-4-0709 | xAI | 47.0 | — | 3 of 4 categories | 51.3 | 48.0 | 35.6 | |
| 355 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 47.0 | $15.00 / $75.00 | 4 of 4 categories | 55.7 | 50.9 | 38.4 | 36.7 |
| 356 | Qwen 2.5qwen/qwen-2.5 | Qwen | 47.0 | — | 1 of 4 categories | 40.6 | |||
| 357 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 47.0 | $3.00 / $15.00 | 4 of 4 categories | 52.0 | 58.8 | 35.7 | 35.1 |
| 358 | Claude Opus 4.8 (medium)anthropic/claude-opus-4.8:medium | Anthropic | 46.9 | $5.00 / $25.00 | 1 of 4 categories | 40.6 | |||
| 359 | Gemini 3.1 Flash Lite (low)google/gemini-3.1-flash-lite:low | 46.9 | $0.25 / $1.50 | 1 of 4 categories | 40.6 | ||||
| 360 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 46.9 | $2.00 / $12.00 | 3 of 4 categories | 43.5 | 29.2 | 61.6 | ||
| 361 | O1 Miniopenai/o1-mini | OpenAI | 46.9 | — | 2 of 4 categories | 45.5 | 41.9 | ||
| 362 | QwQ 32Bqwen/qwq-32b | Qwen | 46.9 | — | 2 of 4 categories | 44.9 | 42.6 | ||
| 363 | Claude Sonnet 4.6 (low)anthropic/claude-sonnet-4.6:low | Anthropic | 46.9 | $3.00 / $15.00 | 1 of 4 categories | 40.5 | |||
| 364 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 46.9 | — | 3 of 4 categories | 45.0 | 48.8 | 40.4 | |
| 365 | Claude Sonnet 4.5 (thinking)anthropic/claude-sonnet-4.5:thinking | Anthropic | 46.9 | $3.00 / $15.00 | 1 of 4 categories | 40.3 | |||
| 366 | Hunyuan TurboStencent/hunyuan-turbos | Tencent | 46.8 | — | 2 of 4 categories | 46.8 | 40.3 | ||
| 367 | Kimi K2 Instructmoonshotai/kimi-k2-instruct | Moonshot AI | 46.8 | — | 2 of 4 categories | 40.8 | 46.2 | ||
| 368 | GPT-5.6 Luna (none)openai/gpt-5.6-luna:none | OpenAI | 46.8 | $0.10 / $0.60 | 1 of 4 categories | 40.1 | |||
| 369 | MiniMax M3 (none)minimax/minimax-m3:none | MiniMax | 46.8 | $0.30 / $1.20 | 1 of 4 categories | 40.1 | |||
| 370 | GLM 4.6z-ai/glm-4.6 | Z.ai | 46.7 | $0.55 / $2.20 | 3 of 4 categories | 45.7 | 44.9 | 42.9 | |
| 371 | Ring Flash 2.0inclusionai/ring-flash-2.0 | inclusionAI | 46.7 | — | 2 of 4 categories | 46.4 | 40.1 | ||
| 372 | Nova 2 Liteamazon/nova-2-lite-v1 | Amazon | 46.7 | $0.30 / $2.50 | 2 of 4 categories | 46.9 | 39.5 | ||
| 373 | O1 Mini (medium)openai/o1-mini:medium | OpenAI | 46.7 | — | 1 of 4 categories | 39.7 | |||
| 374 | Claude 3.7 Sonnet (64K)anthropic/claude-3-7-sonnet:64k | Anthropic | 46.6 | — | 1 of 4 categories | 39.6 | |||
| 375 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Qwen | 46.6 | $0.12 / $0.50 | 2 of 4 categories | 45.3 | 41.0 | ||
| 376 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 46.6 | $0.30 / $1.20 | 3 of 4 categories | 24.4 | 56.0 | 52.3 | |
| 377 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 46.6 | — | 3 of 4 categories | 44.8 | 51.6 | 36.4 | |
| 378 | Gemini 3.1 Flash Lite (minimal)google/gemini-3.1-flash-lite:minimal | 46.6 | $0.25 / $1.50 | 1 of 4 categories | 39.5 | ||||
| 379 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 46.5 | $0.05 / $0.40 | 3 of 4 categories | 39.4 | 43.5 | 49.3 | |
| 380 | Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking | Xiaomi | 46.5 | — | 2 of 4 categories | 41.9 | 43.8 | ||
| 381 | Claude Sonnet 5 (low)anthropic/claude-sonnet-5:low | Anthropic | 46.5 | $2.00 / $10.00 | 1 of 4 categories | 39.2 | |||
| 382 | GPT-5 Nano (minimal)openai/gpt-5-nano:minimal | OpenAI | 46.4 | $0.05 / $0.40 | 1 of 4 categories | 38.8 | |||
| 383 | SWE-1.6 (none)cognition/swe-1.6:none | Cognition | 46.3 | — | 1 of 4 categories | 38.7 | |||
| 384 | Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high | Anthropic | 46.3 | $3.00 / $15.00 | 2 of 4 categories | 35.4 | 49.5 | ||
| 385 | MiniMax M2minimax/minimax-m2 | MiniMax | 46.3 | $0.255 / $1.02 | 3 of 4 categories | 50.5 | 39.2 | 41.3 | |
| 386 | Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct | Qwen | 46.3 | — | 1 of 4 categories | 38.5 | |||
| 387 | Grok Code Fast 1x-ai/grok-code-fast-1 | xAI | 46.2 | — | 1 of 4 categories | 38.5 | |||
| 388 | GPT-5.4 Mini (low)openai/gpt-5.4-mini:low | OpenAI | 46.2 | $0.75 / $4.50 | 1 of 4 categories | 38.3 | |||
| 389 | Command A (03-2025)cohere/command-a-03-2025 | Cohere | 46.1 | — | 2 of 4 categories | 46.0 | 38.2 | ||
| 390 | GLM 5.2 (none)z-ai/glm-5.2:none | Z.ai | 46.1 | $0.462 / $1.452 | 2 of 4 categories | 46.2 | 37.8 | ||
| 391 | Qwen2.5 VL 32B Instructqwen/qwen2.5-vl-32b-instruct | Qwen | 46.1 | — | 1 of 4 categories | 37.9 | |||
| 392 | Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Qwen | 46.0 | $0.065 / $0.26 | 2 of 4 categories | 40.9 | 43.0 | ||
| 393 | Devstral Medium 2507mistralai/devstral-medium-2507 | Mistral AI | 46.0 | — | 1 of 4 categories | 37.9 | |||
| 394 | Mistral Medium 3.5 (none)mistralai/mistral-medium-3-5:none | Mistral | 46.0 | $1.50 / $7.50 | 1 of 4 categories | 37.9 | |||
| 395 | GPT-4oopenai/gpt-4o | OpenAI | 46.0 | $2.50 / $10.00 | 1 of 4 categories | 37.9 | |||
| 396 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | NVIDIA | 46.0 | $0.05 / $0.20 | 2 of 4 categories | 43.2 | 40.6 | ||
| 397 | Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | 46.0 | $0.22 / $0.85 | 2 of 4 categories | 38.7 | 45.1 | ||
| 398 | GPT-5 Miniopenai/gpt-5-mini | OpenAI | 46.0 | $0.25 / $2.00 | 2 of 4 categories | 46.8 | 36.9 | ||
| 399 | gpt-oss-20bopenai/gpt-oss-20b | OpenAI | 46.0 | $0.03 / $0.13 | 2 of 4 categories | 44.0 | 39.6 | ||
| 400 | Claude Sonnet 4 (59K)anthropic/claude-sonnet-4:59k | Anthropic | 46.0 | $3.00 / $15.00 | 1 of 4 categories | 37.6 | |||
| 401 | DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | DeepSeek | 45.9 | $0.27 / $1.12 | 2 of 4 categories | 49.8 | 33.6 | ||
| 402 | GPT-5.4 Mini (none)openai/gpt-5.4-mini:none | OpenAI | 45.9 | $0.75 / $4.50 | 1 of 4 categories | 37.5 | |||
| 403 | Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking | Tencent | 45.9 | — | 2 of 4 categories | 52.0 | 31.3 | ||
| 404 | DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none | DeepSeek | 45.8 | $1.168 / $2.336 | 2 of 4 categories | 41.4 | 41.4 | ||
| 405 | Grok 4.3 (high)x-ai/grok-4.3:high | SpaceXAI | 45.7 | $1.25 / $2.50 | 2 of 4 categories | 30.9 | 51.8 | ||
| 406 | OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct | Allen Institute for AI | 45.6 | — | 2 of 4 categories | 44.4 | 37.6 | ||
| 407 | OLMo 3 32B Thinkallenai/olmo-3-32b-think | Allen Institute for AI | 45.6 | — | 2 of 4 categories | 43.5 | 38.5 | ||
| 408 | Magistral Medium 2506mistralai/magistral-medium-2506 | Mistral AI | 45.5 | — | 2 of 4 categories | 45.7 | 36.0 | ||
| 409 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 45.5 | $0.30 / $2.50 | 4 of 4 categories | 38.6 | 41.4 | 44.1 | 48.4 | |
| 410 | GPT-5.6 Luna (medium)openai/gpt-5.6-luna:medium | OpenAI | 45.3 | $0.10 / $0.60 | 1 of 4 categories | 35.7 | |||
| 411 | Devstral 2mistralai/devstral-2 | Mistral AI | 45.3 | — | 2 of 4 categories | 43.8 | 37.2 | ||
| 412 | Mercury 2inception/mercury-2 | Inception | 45.3 | $0.25 / $0.75 | 1 of 4 categories | 35.6 | |||
| 413 | GLM 4.5z-ai/glm-4.5 | Z.ai | 45.3 | $0.60 / $2.20 | 3 of 4 categories | 44.5 | 48.8 | 32.7 | |
| 414 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 45.1 | — | 4 of 4 categories | 42.3 | 60.3 | 33.9 | 33.9 |
| 415 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 45.1 | — | 3 of 4 categories | 54.7 | 47.3 | 23.0 | |
| 416 | Claude Opus 4.1 (16K)anthropic/claude-opus-4.1:16k | Anthropic | 45.1 | $15.00 / $75.00 | 1 of 4 categories | 34.9 | |||
| 417 | o3 Proopenai/o3-pro | OpenAI | 45.0 | $20.00 / $80.00 | 1 of 4 categories | 34.6 | |||
| 418 | GPT-5.1 (none)openai/gpt-5.1:none | OpenAI | 44.9 | $1.25 / $10.00 | 1 of 4 categories | 34.5 | |||
| 419 | Gemma 3 12Bgoogle/gemma-3-12b-it | 44.8 | $0.05 / $0.15 | 2 of 4 categories | 39.9 | 38.9 | |||
| 420 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 44.7 | — | 3 of 4 categories | 42.5 | 35.7 | 44.9 | ||
| 421 | GPT-4.1openai/gpt-4.1 | OpenAI | 44.6 | $2.00 / $8.00 | 4 of 4 categories | 40.1 | 47.2 | 35.9 | 44.2 |
| 422 | GLM 5.1 (max)z-ai/glm-5.1:max | Z.ai | 44.6 | $0.966 / $3.036 | 1 of 4 categories | 33.4 | |||
| 423 | Claude Opus 4.1 (32K)anthropic/claude-opus-4.1:32k | Anthropic | 44.5 | $15.00 / $75.00 | 1 of 4 categories | 33.2 | |||
| 424 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 44.5 | $1.50 / $7.50 | 4 of 4 categories | 23.8 | 48.1 | 54.8 | 39.8 |
| 425 | Gemini 3.5 Flash Lite (high)google/gemini-3.5-flash-lite:high | 44.4 | $0.30 / $2.50 | 1 of 4 categories | 32.9 | ||||
| 426 | OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think | Allen Institute for AI | 44.3 | — | 2 of 4 categories | 40.3 | 36.8 | ||
| 427 | Mistral Nemomistralai/mistral-nemo | Mistral | 44.3 | $0.019 / $0.03 | 1 of 4 categories | 32.8 | |||
| 428 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 44.3 | $0.10 / $0.40 | 3 of 4 categories | 44.6 | 42.8 | 33.9 | ||
| 429 | Qwen (max)qwen/qwen:max | Qwen | 44.3 | — | 1 of 4 categories | 32.6 | |||
| 430 | GPT-5.6 Luna (low)openai/gpt-5.6-luna:low | OpenAI | 44.3 | $0.10 / $0.60 | 2 of 4 categories | 30.7 | 46.2 | ||
| 431 | GLM 4.6Vz-ai/glm-4.6v | Z.ai | 44.3 | $0.30 / $0.90 | 2 of 4 categories | 49.0 | 27.8 | ||
| 432 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 44.2 | $0.10 / $0.40 | 3 of 4 categories | 47.0 | 42.5 | 31.2 | ||
| 433 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 44.1 | $0.03 / $0.17 | 3 of 4 categories | 37.9 | 37.8 | 44.6 | |
| 434 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 44.0 | — | 3 of 4 categories | 40.7 | 39.3 | 39.6 | ||
| 435 | Grok 2 1212x-ai/grok-2-1212 | xAI | 43.9 | — | 1 of 4 categories | 31.5 | |||
| 436 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 43.9 | $1.00 / $5.00 | 3 of 4 categories | 33.4 | 43.4 | 42.3 | |
| 437 | DeepSeek V3deepseek/deepseek-chat | DeepSeek | 43.7 | $0.2574 / $1.0287 | 2 of 4 categories | 41.1 | 33.2 | ||
| 438 | Solar Pro 4upstage/solar-pro4 | Upstage | 43.6 | $0.03 / $0.12 | 2 of 4 categories | 24.9 | 49.3 | ||
| 439 | Gemma 3n E4Bgoogle/gemma-3n-e4b-it | 43.6 | — | 2 of 4 categories | 39.4 | 34.6 | |||
| 440 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 43.5 | $0.40 / $1.60 | 4 of 4 categories | 37.1 | 41.1 | 42.4 | 40.2 |
| 441 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Qwen | 43.4 | $0.36 / $0.40 | 2 of 4 categories | 42.3 | 31.1 | ||
| 442 | Claude Opus 4.7 (medium)anthropic/claude-opus-4.7:medium | Anthropic | 43.3 | $5.00 / $25.00 | 1 of 4 categories | 29.6 | |||
| 443 | Gemma 3 4Bgoogle/gemma-3-4b-it | 43.2 | — | 2 of 4 categories | 38.3 | 34.3 | |||
| 444 | Magistral Small 2506mistralai/magistral-small-2506 | Mistral AI | 43.2 | — | 1 of 4 categories | 29.3 | |||
| 445 | Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | 43.2 | $2.50 / $10.00 | 2 of 4 categories | 38.7 | 33.8 | ||
| 446 | GLM 5.2 (high)z-ai/glm-5.2:high | Z.ai | 43.1 | $0.462 / $1.452 | 1 of 4 categories | 29.2 | |||
| 447 | SWE Llamaprinceton-nlp/swe-llama | Princeton NLP | 43.1 | — | 1 of 4 categories | 29.0 | |||
| 448 | Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0 | Amazon | 43.1 | — | 2 of 4 categories | 39.0 | 33.2 | ||
| 449 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 43.1 | $0.60 / $1.80 | 3 of 4 categories | 47.6 | 41.6 | 25.9 | |
| 450 | Grok Build 0.1x-ai/grok-build-0.1 | SpaceXAI | 43.1 | $1.00 / $2.00 | 1 of 4 categories | 28.9 | |||
| 451 | OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct | Allen Institute for AI | 43.0 | — | 2 of 4 categories | 38.5 | 33.3 | ||
| 452 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 43.0 | $0.05 / $0.40 | 3 of 4 categories | 44.7 | 43.9 | 26.1 | |
| 453 | Command R (08-2024)cohere/command-r-08-2024 | Cohere | 43.0 | $0.15 / $0.60 | 2 of 4 categories | 38.8 | 32.9 | ||
| 454 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 43.0 | $5.00 / $15.00 | 3 of 4 categories | 41.8 | 29.2 | 43.5 | |
| 455 | Step 3stepfun/step-3 | StepFun | 43.0 | — | 3 of 4 categories | 48.0 | 42.0 | 24.4 | |
| 456 | Granite 4.1 8Bibm-granite/granite-4.1-8b | IBM | 42.9 | $0.05 / $0.10 | 2 of 4 categories | 32.3 | 39.0 | ||
| 457 | QwQ 32B Previewqwen/qwq-32b-preview | Qwen | 42.8 | — | 2 of 4 categories | 37.9 | 33.0 | ||
| 458 | GPT 4 1106 Previewopenai/gpt-4-1106-preview | OpenAI | 42.6 | — | 2 of 4 categories | 39.1 | 31.3 | ||
| 459 | SWE Llama 13Bprinceton-nlp/swe-llama-13b | Princeton NLP | 42.5 | — | 1 of 4 categories | 27.4 | |||
| 460 | Mistral Large 2407mistralai/mistral-large-2407 | Mistral | 42.4 | $2.00 / $6.00 | 2 of 4 categories | 42.0 | 27.4 | ||
| 461 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 42.3 | — | 3 of 4 categories | 40.9 | 34.9 | 35.4 | |
| 462 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 42.2 | $0.10 / $0.40 | 3 of 4 categories | 44.3 | 29.8 | 36.5 | |
| 463 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Meta | 41.8 | $0.10 / $0.32 | 2 of 4 categories | 41.0 | 25.9 | ||
| 464 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 41.7 | — | 3 of 4 categories | 39.2 | 34.0 | 35.1 | |
| 465 | Mistral Large 2411mistralai/mistral-large-2411 | Mistral AI | 41.6 | — | 2 of 4 categories | 41.1 | 25.0 | ||
| 466 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 41.5 | — | 3 of 4 categories | 48.4 | 39.8 | 19.0 | |
| 467 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 41.4 | $2.50 / $10.00 | 3 of 4 categories | 36.4 | 43.2 | 27.2 | |
| 468 | Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | 41.4 | $2.00 / $6.00 | 2 of 4 categories | 38.4 | 26.9 | ||
| 469 | Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501 | Mistral AI | 41.3 | — | 2 of 4 categories | 39.6 | 25.3 | ||
| 470 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 41.3 | — | 3 of 4 categories | 41.3 | 26.7 | 38.2 | ||
| 471 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 41.3 | — | 3 of 4 categories | 39.8 | 26.1 | 40.2 | ||
| 472 | GPT 3.5openai/gpt-3.5 | OpenAI | 41.0 | — | 1 of 4 categories | 22.7 | |||
| 473 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 41.0 | — | 3 of 4 categories | 48.8 | 26.7 | 29.2 | |
| 474 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Meta | 41.0 | $0.40 / $0.40 | 2 of 4 categories | 40.2 | 23.4 | ||
| 475 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 40.9 | — | 3 of 4 categories | 50.9 | 33.6 | 20.0 | |
| 476 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 40.9 | — | 3 of 4 categories | 42.4 | 35.8 | 26.0 | |
| 477 | Molmo 2 8Ballenai/molmo-2-8b | Allen Institute for AI | 40.8 | — | 1 of 4 categories | 22.3 | |||
| 478 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 40.6 | — | 3 of 4 categories | 41.7 | 37.2 | 23.9 | |
| 479 | Gemma 3 27Bgoogle/gemma-3-27b-it | 40.5 | $0.08 / $0.45 | 3 of 4 categories | 42.7 | 35.8 | 24.0 | ||
| 480 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 40.2 | $2.50 / $10.00 | 3 of 4 categories | 36.9 | 26.4 | 37.6 | |
| 481 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 40.1 | — | 3 of 4 categories | 38.1 | 26.0 | 36.3 | ||
| 482 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 40.1 | — | 3 of 4 categories | 33.4 | 31.6 | 35.2 | |
| 483 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 39.9 | $10.00 / $30.00 | 3 of 4 categories | 34.1 | 27.9 | 37.4 | |
| 484 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Meta | 39.8 | $0.05 / $0.08 | 2 of 4 categories | 38.0 | 21.0 | ||
| 485 | Claude 2anthropic/claude-2 | Anthropic | 39.8 | — | 2 of 4 categories | 33.7 | 25.1 | ||
| 486 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 39.8 | $0.15 / $0.60 | 3 of 4 categories | 33.2 | 28.5 | 36.8 | |
| 487 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 39.7 | — | 3 of 4 categories | 39.5 | 22.8 | 35.7 | ||
| 488 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 39.5 | — | 3 of 4 categories | 53.5 | 23.2 | 20.7 | |
| 489 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 39.5 | — | 3 of 4 categories | 35.0 | 26.2 | 36.0 | |
| 490 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 39.3 | — | 3 of 4 categories | 40.1 | 21.5 | 34.8 | |
| 491 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 38.8 | — | 4 of 4 categories | 34.9 | 34.4 | 35.7 | 27.7 | |
| 492 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 37.7 | — | 4 of 4 categories | 35.7 | 35.4 | 31.1 | 23.6 |
| 493 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 37.6 | — | 3 of 4 categories | 42.9 | 26.9 | 17.9 | |
| 494 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 36.7 | $0.25 / $1.25 | 3 of 4 categories | 27.8 | 20.9 | 34.6 | |
| 495 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 36.0 | — | 4 of 4 categories | 34.2 | 33.7 | 26.5 | 21.1 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model ranks here on however many of the 4 text categories it has been scored in, and the coverage column says how many that is. Each category score is itself a shrunk mean of that model's benchmark percentiles, so a composite standing on one category sits near the middle of the board rather than at the top of it.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- DeepSWE v1.1
Pass@1 across all 113 DeepSWE v1.1 tasks, run by Datacurve with mini-swe-agent through Pier; reasoning-effort configurations remain separate.
- Epoch AI Benchmarking HubCC BY 4.0
Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- FrontierCode 1.1
Main weighted rubric Scores on FrontierCode 1.1's 100 hardest tasks, run by Cognition across five trials; results identify model, harness, and effort.
- FrontierSWE
Mean@5 Dominance across FrontierSWE's complete 17-task cohort; results identify the evaluated model and harness.
- LiveCodeBenchMIT
Maintainer-published Code Generation pass@1 over the 454-problem release_v6 window (2024-08-01 through 2025-05-01), from the stable 2025-08-01 snapshot; it does not cover newer frontier families.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.
- MCP AtlasMIT
All-1,000-task Pass Rate from Scale's current Performance Comparison, using the standardized MCP loop and 100-tool-call budget.
- Open ASR Leaderboard
Word error rates from the Hugging Face Open ASR Leaderboard.
- SWE-bench
Resolve rates published by the SWE-bench maintainers.
- Terminal-Bench 2.1Apache-2.0
Published Accuracy over 89 Terminal-Bench 2.1 tasks, run and verified by the benchmark team; results are model plus agent harness plus effort, not model-only evaluations.
- TTS Arena V2
Elo ratings from TTS Arena V2.
- Warden
Security review results published by Warden.