Best LLM for coding
Issue resolution, code generation, private mergeability tasks, cohort-relative engineering performance, webdev preference and security review contribute once each. LiveCodeBench is the stable 2025-08-01 snapshot of 454 problems dated 2024-08-01 through 2025-05-01, so it does not cover newer frontier families. Every row prints coverage and marks partial evidence.
Aldena runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | % | % | % | % | % | % | % | % | % | % | elo | elo | % | count |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 86.1 | $5.00 / $25.00 | 5 of 15 benchmarks | 73.7% | 48.0% | 50.0% | 1530 | 1692 | ||||||||||
| 2 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 82.2 | $5.00 / $25.00 | 4 of 15 benchmarks | 72.8% | 48.0% | 1531 | 1663 | |||||||||||
| 3 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 81.8 | $10.00 / $50.00 | 3 of 15 benchmarks | 88.2% | 1554 | 1627 | ||||||||||||
| 4 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 80.4 | $5.00 / $30.00 | 4 of 15 benchmarks | 70.7% | 46.8% | 1526 | 1622 | |||||||||||
| 5 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 79.1 | $3.00 / $15.00 | 5 of 15 benchmarks | 68.5% | 1543 | 1674 | 37.2% | 44.0 | ||||||||||
| 6 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 78.8 | $10.00 / $50.00 | 3 of 15 benchmarks | 68.6% | 52.7% | 63.9% | ||||||||||||
| 7 | Claude Fable 5 (max)anthropic/claude-fable-5:max | Anthropic | 78.7 | $10.00 / $50.00 | 3 of 15 benchmarks | 69.7% | 51.6% | 39.5% | ||||||||||||
| 8 | Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | 77.3 | $5.00 / $25.00 | 5 of 15 benchmarks | 76.7% | 52.6% | 70.7% | 1523 | 1468 | ||||||||||
| 9 | GPT-5.6 Sol (max)openai/gpt-5.6-sol:max | OpenAI | 77.0 | $5.00 / $30.00 | 3 of 15 benchmarks | 72.7% | 47.5% | 39.0% | ||||||||||||
| 10 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 76.7 | $2.00 / $6.00 | 2 of 15 benchmarks | 1529 | 1667 | |||||||||||||
| 11 | Claude Fable 5 (xhigh)anthropic/claude-fable-5:xhigh | Anthropic | 76.3 | $10.00 / $50.00 | 2 of 15 benchmarks | 69.9% | 53.5% | |||||||||||||
| 12 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 75.8 | $5.00 / $25.00 | 5 of 15 benchmarks | 78.7% | 72.0% | 48.9% | 1547 | 1537 | ||||||||||
| 13 | Grok 4.6 (high)x-ai/grok-4.6:high | SpaceXAI | 75.5 | $2.00 / $6.00 | 4 of 15 benchmarks | 65.2% | 48.0% | 1512 | 1631 | |||||||||||
| 14 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 74.6 | $2.00 / $6.00 | 3 of 15 benchmarks | 72.1% | 1521 | 1555 | ||||||||||||
| 15 | Claude Opus 5 (medium)anthropic/claude-opus-5:medium | Anthropic | 74.3 | $5.00 / $25.00 | 2 of 15 benchmarks | 68.9% | 53.4% | |||||||||||||
| 16 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 73.5 | $1.25 / $4.25 | 2 of 15 benchmarks | 1531 | 1538 | |||||||||||||
| 17 | Claude Opus 5 (xhigh)anthropic/claude-opus-5:xhigh | Anthropic | 73.0 | $5.00 / $25.00 | 2 of 15 benchmarks | 73.2% | 43.6% | |||||||||||||
| 18 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 72.2 | $0.50 / $3.00 | 4 of 15 benchmarks | 75.4% | 72.7% | 1508 | 1438 | ||||||||||||
| 19 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 72.0 | $5.00 / $25.00 | 3 of 15 benchmarks | 66.5% | 1524 | 1539 | ||||||||||||
| 20 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 71.5 | $5.00 / $25.00 | 3 of 15 benchmarks | 56.3% | 1547 | 1558 | ||||||||||||
| 21 | o4 Mini Highopenai/o4-mini-high | OpenAI | 71.2 | $1.10 / $4.40 | 1 of 15 benchmarks | 80.2% | ||||||||||||||
| 22 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 71.0 | $5.00 / $25.00 | 4 of 15 benchmarks | 35.2% | 31.1% | 1552 | 1557 | |||||||||||
| 23 | GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh | OpenAI | 70.8 | — | 1 of 15 benchmarks | 80.6% | ||||||||||||||
| 24 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 70.5 | $5.00 / $30.00 | 4 of 15 benchmarks | 67.0% | 43.0% | 34.3% | 1507 | |||||||||||
| 25 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 70.2 | $3.00 / $15.00 | 5 of 15 benchmarks | 75.2% | 1528 | 1524 | 29.1% | 32.0 | ||||||||||
| 26 | Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k | Anthropic | 70.0 | $5.00 / $25.00 | 2 of 15 benchmarks | 1530 | 1494 | |||||||||||||
| 27 | Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium | Anthropic | 70.0 | $5.00 / $25.00 | 1 of 15 benchmarks | 79.2% | ||||||||||||||
| 28 | Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | 70.0 | $1.027 / $6.162 | 3 of 15 benchmarks | 76.7% | 1509 | 1479 | ||||||||||||
| 29 | Claude Fable 5 (medium)anthropic/claude-fable-5:medium | Anthropic | 69.9 | $10.00 / $50.00 | 2 of 15 benchmarks | 65.4% | 49.8% | |||||||||||||
| 30 | o3 (high)openai/o3:high | OpenAI | 69.9 | $2.00 / $8.00 | 1 of 15 benchmarks | 75.8% | ||||||||||||||
| 31 | Gemini 3.7 Flash (high)google/gemini-3.7-flash:high | 69.8 | $0.375 / $1.875 | 3 of 15 benchmarks | 65.3% | 42.2% | 1587 | |||||||||||||
| 32 | Doubao-Seed-Codebytedance/doubao-seed-code | ByteDance | 69.7 | — | 1 of 15 benchmarks | 78.8% | ||||||||||||||
| 33 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 69.2 | $5.00 / $30.00 | 7 of 15 benchmarks | 64.4% | 40.6% | 10.0% | 1520 | 1487 | 47.7% | 72.0 | ||||||||
| 34 | Muse Sparkmeta/muse-spark | Meta | 69.1 | — | 1 of 15 benchmarks | 1526 | ||||||||||||||
| 35 | GLM-5.3z-ai/glm-5.3 | Z.ai | 69.1 | — | 1 of 15 benchmarks | 78.1% | ||||||||||||||
| 36 | Muse Spark 1.2 (xhigh)meta/muse-spark-1.2:xhigh | Meta | 68.9 | $1.25 / $4.25 | 3 of 15 benchmarks | 54.9% | 1533 | 1535 | ||||||||||||
| 37 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 68.9 | $5.00 / $25.00 | 2 of 15 benchmarks | 46.5% | 28.6% | |||||||||||||
| 38 | DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high | DeepSeek | 68.8 | $1.168 / $2.336 | 2 of 15 benchmarks | 1489 | 1584 | |||||||||||||
| 39 | GPT-5.6 Sol (high)openai/gpt-5.6-sol:high | OpenAI | 68.6 | $5.00 / $30.00 | 3 of 15 benchmarks | 69.4% | 45.1% | 20.0% | ||||||||||||
| 40 | o4 Mini (medium)openai/o4-mini:medium | OpenAI | 68.5 | $1.10 / $4.40 | 1 of 15 benchmarks | 74.2% | ||||||||||||||
| 41 | Hy3tencent/hy3 | Tencent | 67.8 | $0.132 / $0.528 | 2 of 15 benchmarks | 1503 | 1522 | |||||||||||||
| 42 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 67.6 | $0.0643 / $0.1285 | 2 of 15 benchmarks | 1480 | 1581 | |||||||||||||
| 43 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 67.6 | $0.435 / $0.87 | 2 of 15 benchmarks | 1520 | 1474 | |||||||||||||
| 44 | DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max | DeepSeek | 67.5 | $1.168 / $2.336 | 2 of 15 benchmarks | 77.6% | 62.8% | |||||||||||||
| 45 | Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high | Anthropic | 67.4 | $5.00 / $25.00 | 1 of 15 benchmarks | 76.8% | ||||||||||||||
| 46 | GPT-5.6 Terra (max)openai/gpt-5.6-terra:max | OpenAI | 67.2 | $1.00 / $6.00 | 2 of 15 benchmarks | 69.6% | 41.3% | |||||||||||||
| 47 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 67.1 | $1.75 / $14.00 | 1 of 15 benchmarks | 1515 | ||||||||||||||
| 48 | Grok 4.6x-ai/grok-4.6 | SpaceXAI | 67.0 | $2.00 / $6.00 | 1 of 15 benchmarks | 77.9% | ||||||||||||||
| 49 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 66.8 | $1.475 / $4.425 | 4 of 15 benchmarks | 77.3% | 9.5% | 1525 | 1517 | |||||||||||
| 50 | Ernie 5.1baidu/ernie-5.1 | Baidu | 66.7 | — | 1 of 15 benchmarks | 1514 | ||||||||||||||
| 51 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 66.7 | — | 1 of 15 benchmarks | 1514 | ||||||||||||||
| 52 | Qwen3.5 Max Previewqwen/qwen3.5-max-preview | Qwen | 66.4 | — | 1 of 15 benchmarks | 1513 | ||||||||||||||
| 53 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 66.4 | $2.50 / $15.00 | 4 of 15 benchmarks | 76.9% | 15.6% | 1521 | 1463 | |||||||||||
| 54 | Gemini 3.7 Flash (medium)google/gemini-3.7-flash:medium | 66.3 | $0.375 / $1.875 | 2 of 15 benchmarks | 65.5% | 43.6% | ||||||||||||||
| 55 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 66.2 | $1.00 / $6.00 | 4 of 15 benchmarks | 60.2% | 38.8% | 1516 | 1521 | |||||||||||
| 56 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 66.1 | — | 1 of 15 benchmarks | 1513 | ||||||||||||||
| 57 | Claude Fable 5 (low)anthropic/claude-fable-5:low | Anthropic | 66.1 | $10.00 / $50.00 | 2 of 15 benchmarks | 59.6% | 48.0% | |||||||||||||
| 58 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 66.0 | $1.50 / $9.00 | 2 of 15 benchmarks | 1507 | 1489 | ||||||||||||||
| 59 | Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k | Anthropic | 66.0 | $15.00 / $75.00 | 1 of 15 benchmarks | 1512 | ||||||||||||||
| 60 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 65.7 | $2.00 / $6.00 | 4 of 15 benchmarks | 53.8% | 42.4% | 38.4% | 41.0 | |||||||||||
| 61 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 65.7 | $0.50 / $3.00 | 1 of 15 benchmarks | 75.8% | |||||||||||||||
| 62 | MiniMax M2.5 (high)minimax/minimax-m2.5:high | MiniMax | 65.7 | $0.22 / $0.90 | 1 of 15 benchmarks | 75.8% | ||||||||||||||
| 63 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 65.5 | $5.00 / $25.00 | 3 of 15 benchmarks | 83.5% | 38.5% | 19.1% | ||||||||||||
| 64 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 65.1 | — | 1 of 15 benchmarks | 1508 | ||||||||||||||
| 65 | Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | 65.1 | $2.00 / $12.00 | 1 of 15 benchmarks | 75.6% | |||||||||||||||
| 66 | Kimi K3 (none)moonshotai/kimi-k3:none | MoonshotAI | 65.0 | $3.00 / $15.00 | 1 of 15 benchmarks | 44.2% | ||||||||||||||
| 67 | Gemini 3 Progoogle/gemini-3-pro | 65.0 | — | 3 of 15 benchmarks | 68.7% | 1518 | 1438 | |||||||||||||
| 68 | MiniMax M3minimax/minimax-m3 | MiniMax | 64.8 | $0.30 / $1.20 | 2 of 15 benchmarks | 1497 | 1490 | |||||||||||||
| 69 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 64.7 | $0.32 / $1.28 | 1 of 15 benchmarks | 1506 | ||||||||||||||
| 70 | EXAONE 4.0 32Blg-ai/exaone-4.0-32b | LG AI Research | 64.5 | — | 1 of 15 benchmarks | 70.0% | ||||||||||||||
| 71 | Grok 4.6 (medium)x-ai/grok-4.6:medium | SpaceXAI | 64.4 | $2.00 / $6.00 | 1 of 15 benchmarks | 67.5% | ||||||||||||||
| 72 | GPT-5.5openai/gpt-5.5 | OpenAI | 64.3 | $5.00 / $30.00 | 3 of 15 benchmarks | 64.5% | 1510 | 1457 | ||||||||||||
| 73 | GLM 5z-ai/glm-5 | Z.ai | 64.1 | $0.60 / $1.92 | 4 of 15 benchmarks | 72.1% | 69.7% | 1497 | 1436 | |||||||||||
| 74 | GPT-5.3-Codex (high)openai/gpt-5.3-codex:high | OpenAI | 64.0 | $1.75 / $14.00 | 1 of 15 benchmarks | 74.8% | ||||||||||||||
| 75 | Seed 2.1 Pro Previewbytedance/seed-2.1-pro-preview | ByteDance | 64.0 | — | 1 of 15 benchmarks | 1522 | ||||||||||||||
| 76 | R1 0528deepseek/deepseek-r1-0528 | DeepSeek | 63.7 | $0.50 / $2.15 | 2 of 15 benchmarks | 73.1% | 1464 | |||||||||||||
| 77 | GPT-5openai/gpt-5 | OpenAI | 63.6 | $1.25 / $10.00 | 1 of 15 benchmarks | 74.4% | ||||||||||||||
| 78 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 63.5 | $2.00 / $10.00 | 4 of 15 benchmarks | 48.2% | 39.4% | 1523 | 1540 | |||||||||||
| 79 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 63.5 | $15.00 / $75.00 | 1 of 15 benchmarks | 1499 | ||||||||||||||
| 80 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 63.4 | $0.30 / $2.50 | 2 of 15 benchmarks | 1503 | 1449 | ||||||||||||||
| 81 | OpenReasoning Nemotron 32Bnvidia/openreasoning-nemotron-32b | NVIDIA | 63.2 | — | 1 of 15 benchmarks | 69.8% | ||||||||||||||
| 82 | GPT-5.6 Luna (max)openai/gpt-5.6-luna:max | OpenAI | 63.0 | $0.10 / $0.60 | 2 of 15 benchmarks | 67.2% | 39.8% | |||||||||||||
| 83 | GLM 5.1z-ai/glm-5.1 | Z.ai | 62.9 | $0.966 / $3.036 | 4 of 15 benchmarks | 74.2% | 25.7% | 1515 | 1510 | |||||||||||
| 84 | GLM 5.2z-ai/glm-5.2 | Z.ai | 62.9 | $0.462 / $1.452 | 1 of 15 benchmarks | 67.5% | ||||||||||||||
| 85 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 62.8 | $0.10 / $0.60 | 4 of 15 benchmarks | 56.9% | 38.9% | 1499 | 1517 | |||||||||||
| 86 | Grok 4.6 (xhigh)x-ai/grok-4.6:xhigh | SpaceXAI | 62.7 | $2.00 / $6.00 | 1 of 15 benchmarks | 66.7% | ||||||||||||||
| 87 | GPT-5.3 Chatopenai/gpt-5.3-chat | OpenAI | 62.5 | — | 1 of 15 benchmarks | 1496 | ||||||||||||||
| 88 | Claude Opus 4.8 (xhigh)anthropic/claude-opus-4.8:xhigh | Anthropic | 62.2 | $5.00 / $25.00 | 2 of 15 benchmarks | 54.4% | 45.5% | |||||||||||||
| 89 | Grok 4.1x-ai/grok-4.1 | xAI | 62.1 | — | 1 of 15 benchmarks | 1492 | ||||||||||||||
| 90 | GPT-5.6 Luna (high)openai/gpt-5.6-luna:high | OpenAI | 62.1 | $0.10 / $0.60 | 4 of 15 benchmarks | 44.3% | 35.9% | 41.9% | 57.0 | |||||||||||
| 91 | SWE-1.7 (none)cognition/swe-1.7:none | Cognition | 61.5 | — | 1 of 15 benchmarks | 42.0% | ||||||||||||||
| 92 | Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k | Anthropic | 61.5 | $3.00 / $15.00 | 2 of 15 benchmarks | 1519 | 1392 | |||||||||||||
| 93 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 61.5 | $0.57 / $2.85 | 2 of 15 benchmarks | 1502 | 1436 | |||||||||||||
| 94 | o3openai/o3 | OpenAI | 61.5 | $2.00 / $8.00 | 3 of 15 benchmarks | 58.4% | 36.0% | 1460 | ||||||||||||
| 95 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 61.4 | $0.462 / $1.452 | 5 of 15 benchmarks | 78.7% | 43.8% | 9.5% | 1506 | 1585 | ||||||||||
| 96 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 61.3 | — | 1 of 15 benchmarks | 72.9% | |||||||||||||||
| 97 | Mimo v2 Proxiaomi/mimo-v2-pro | Xiaomi | 61.2 | — | 2 of 15 benchmarks | 1503 | 1434 | |||||||||||||
| 98 | Ernie 5.0 0110baidu/ernie-5.0-0110 | Baidu | 61.2 | — | 1 of 15 benchmarks | 1490 | ||||||||||||||
| 99 | GLM 5 (high)z-ai/glm-5:high | Z.ai | 60.8 | $0.60 / $1.92 | 1 of 15 benchmarks | 72.8% | ||||||||||||||
| 100 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 60.8 | — | 1 of 15 benchmarks | 1486 | ||||||||||||||
| 101 | Grok 4.5 (medium)x-ai/grok-4.5:medium | SpaceXAI | 60.7 | $2.00 / $6.00 | 1 of 15 benchmarks | 41.9% | ||||||||||||||
| 102 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 60.6 | $0.14 / $0.28 | 2 of 15 benchmarks | 1491 | 1438 | |||||||||||||
| 103 | OpenCodeReasoning Nemotron 1.1 32Bnvidia/opencodereasoning-nemotron-1.1-32b | NVIDIA | 60.5 | — | 1 of 15 benchmarks | 66.8% | ||||||||||||||
| 104 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 60.4 | — | 2 of 15 benchmarks | 1505 | 1405 | |||||||||||||
| 105 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 60.4 | $0.0643 / $0.1285 | 1 of 15 benchmarks | 1483 | ||||||||||||||
| 106 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 60.3 | — | 5 of 15 benchmarks | 61.0% | 51.7% | 33.8% | 31.3% | 1430 | ||||||||||
| 107 | Claude Opus 5 (low)anthropic/claude-opus-5:low | Anthropic | 60.2 | $5.00 / $25.00 | 2 of 15 benchmarks | 58.1% | 42.0% | |||||||||||||
| 108 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 59.9 | $1.25 / $10.00 | 2 of 15 benchmarks | 68.0% | 1491 | |||||||||||||
| 109 | GPT-5.6 Sol (medium)openai/gpt-5.6-sol:medium | OpenAI | 59.5 | $5.00 / $30.00 | 2 of 15 benchmarks | 61.1% | 39.9% | |||||||||||||
| 110 | Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high | Anthropic | 59.5 | $3.00 / $15.00 | 1 of 15 benchmarks | 71.4% | ||||||||||||||
| 111 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 59.4 | $1.75 / $14.00 | 1 of 15 benchmarks | 23.0% | ||||||||||||||
| 112 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 59.3 | $0.60 / $3.60 | 1 of 15 benchmarks | 1475 | ||||||||||||||
| 113 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 59.1 | $0.75 / $3.75 | 4 of 15 benchmarks | 46.7% | 33.7% | 1522 | 1537 | ||||||||||||
| 114 | DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking | DeepSeek | 59.0 | $0.27 / $0.41 | 1 of 15 benchmarks | 1475 | ||||||||||||||
| 115 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 58.8 | $2.00 / $12.00 | 4 of 15 benchmarks | 34.4% | 14.3% | 1521 | 1447 | ||||||||||||
| 116 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 58.8 | $3.00 / $15.00 | 5 of 15 benchmarks | 57.0% | 58.3% | 35.6% | 47.1% | 1449 | ||||||||||
| 117 | Kimi K2 0905moonshotai/kimi-k2-0905 | MoonshotAI | 58.7 | $0.60 / $2.50 | 2 of 15 benchmarks | 71.2% | 1468 | |||||||||||||
| 118 | LongCat Flash Chatmeituan/longcat-flash-chat | Meituan | 58.7 | — | 1 of 15 benchmarks | 1474 | ||||||||||||||
| 119 | Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max | Anthropic | 58.7 | $2.00 / $10.00 | 2 of 15 benchmarks | 53.9% | 42.4% | |||||||||||||
| 120 | GLM 4.7z-ai/glm-4.7 | Z.ai | 58.7 | $0.40 / $1.75 | 2 of 15 benchmarks | 1485 | 1434 | |||||||||||||
| 121 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 58.6 | $3.00 / $15.00 | 1 of 15 benchmarks | 1473 | ||||||||||||||
| 122 | Qwen3 Maxqwen/qwen3-max | Qwen | 58.4 | $0.78 / $3.90 | 1 of 15 benchmarks | 1473 | ||||||||||||||
| 123 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 58.3 | $1.75 / $14.00 | 3 of 15 benchmarks | 73.8% | 66.7% | 1490 | ||||||||||||
| 124 | Kimi K2.5 (high)moonshotai/kimi-k2.5:high | MoonshotAI | 58.3 | $0.57 / $2.85 | 1 of 15 benchmarks | 70.8% | ||||||||||||||
| 125 | Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | Qwen | 58.3 | $0.09 / $0.55 | 1 of 15 benchmarks | 1472 | ||||||||||||||
| 126 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 58.3 | — | 2 of 15 benchmarks | 1511 | 1374 | |||||||||||||
| 127 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 58.2 | $0.39 / $2.34 | 2 of 15 benchmarks | 1491 | 1400 | |||||||||||||
| 128 | Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203 | Baidu | 58.2 | — | 1 of 15 benchmarks | 1472 | ||||||||||||||
| 129 | Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high | Anthropic | 57.9 | $5.00 / $25.00 | 5 of 15 benchmarks | 26.6% | 1552 | 1545 | 26.7% | 24.0 | ||||||||||
| 130 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 57.8 | — | 1 of 15 benchmarks | 1468 | ||||||||||||||
| 131 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 57.7 | $1.25 / $10.00 | 2 of 15 benchmarks | 71.5% | 1419 | |||||||||||||
| 132 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 57.6 | $5.00 / $25.00 | 6 of 15 benchmarks | 51.8% | 41.0% | 1533 | 1564 | 24.4% | 24.0 | |||||||||
| 133 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 57.6 | $1.20 / $4.00 | 2 of 15 benchmarks | 1490 | 1400 | |||||||||||||
| 134 | DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high | DeepSeek | 57.6 | $0.269 / $0.40 | 1 of 15 benchmarks | 70.0% | ||||||||||||||
| 135 | GPT-5.4 (medium)openai/gpt-5.4:medium | OpenAI | 57.4 | $2.50 / $15.00 | 1 of 15 benchmarks | 1442 | ||||||||||||||
| 136 | GPT-5.2openai/gpt-5.2 | OpenAI | 57.4 | $1.75 / $14.00 | 3 of 15 benchmarks | 69.0% | 1482 | 1418 | ||||||||||||
| 137 | o4 Mini (low)openai/o4-mini:low | OpenAI | 57.2 | $1.10 / $4.40 | 1 of 15 benchmarks | 65.9% | ||||||||||||||
| 138 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 57.1 | $0.26 / $1.04 | 1 of 15 benchmarks | 1465 | ||||||||||||||
| 139 | Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high | 57.0 | — | 1 of 15 benchmarks | 69.6% | |||||||||||||||
| 140 | GPT-5.4openai/gpt-5.4 | OpenAI | 57.0 | $2.50 / $15.00 | 3 of 15 benchmarks | 46.0% | 1514 | 1390 | ||||||||||||
| 141 | GPT-5 (high)openai/gpt-5:high | OpenAI | 56.9 | $1.25 / $10.00 | 3 of 15 benchmarks | 73.5% | 12.7% | 1469 | ||||||||||||
| 142 | DeepSeek V3.1 Terminus (thinking)deepseek/deepseek-v3.1-terminus:thinking | DeepSeek | 56.6 | $0.27 / $0.95 | 1 of 15 benchmarks | 1463 | ||||||||||||||
| 143 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 56.5 | — | 1 of 15 benchmarks | 1462 | ||||||||||||||
| 144 | Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh | Qwen | 56.5 | $2.00 / $6.00 | 1 of 15 benchmarks | 57.5% | ||||||||||||||
| 145 | o3 Mini Highopenai/o3-mini-high | OpenAI | 56.4 | $1.10 / $4.40 | 2 of 15 benchmarks | 67.4% | 1435 | |||||||||||||
| 146 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 56.1 | $1.50 / $9.00 | 5 of 15 benchmarks | 79.3% | 36.1% | 4.8% | 1509 | 1499 | |||||||||||
| 147 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 56.0 | $0.30 / $1.20 | 2 of 15 benchmarks | 1479 | 1397 | |||||||||||||
| 148 | Gemma 4 31Bgoogle/gemma-4-31b-it | 56.0 | $0.10 / $0.34 | 2 of 15 benchmarks | 1499 | 1364 | ||||||||||||||
| 149 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 56.0 | $0.20 / $1.25 | 1 of 15 benchmarks | 1460 | ||||||||||||||
| 150 | Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh | Anthropic | 56.0 | $2.00 / $10.00 | 2 of 15 benchmarks | 49.7% | 42.7% | |||||||||||||
| 151 | Grok 4.5 (low)x-ai/grok-4.5:low | SpaceXAI | 55.8 | $2.00 / $6.00 | 1 of 15 benchmarks | 37.9% | ||||||||||||||
| 152 | Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high | Anthropic | 55.7 | $1.00 / $5.00 | 1 of 15 benchmarks | 66.6% | ||||||||||||||
| 153 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 55.5 | — | 1 of 15 benchmarks | 1459 | ||||||||||||||
| 154 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 55.4 | $0.50 / $3.00 | 2 of 15 benchmarks | 1491 | 1383 | ||||||||||||||
| 155 | XBai o4 (medium)metastone/xbai-o4:medium | MetaStoneTec | 55.2 | — | 1 of 15 benchmarks | 65.0% | ||||||||||||||
| 156 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 55.1 | $2.50 / $15.00 | 2 of 15 benchmarks | 51.8% | 25.4% | |||||||||||||
| 157 | GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium | OpenAI | 55.1 | $1.25 / $10.00 | 1 of 15 benchmarks | 66.0% | ||||||||||||||
| 158 | DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking | DeepSeek | 55.0 | $0.25 / $0.95 | 1 of 15 benchmarks | 1457 | ||||||||||||||
| 159 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 54.7 | — | 1 of 15 benchmarks | 1455 | ||||||||||||||
| 160 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 54.6 | $0.40 / $4.00 | 1 of 15 benchmarks | 1455 | ||||||||||||||
| 161 | Kimi K2 0711moonshotai/kimi-k2 | MoonshotAI | 54.6 | $0.57 / $2.30 | 2 of 15 benchmarks | 65.4% | 1461 | |||||||||||||
| 162 | Claude Opus 4.5 (128K)anthropic/claude-opus-4.5:128k | Anthropic | 54.5 | $5.00 / $25.00 | 1 of 15 benchmarks | 14.3% | ||||||||||||||
| 163 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 54.5 | $3.00 / $15.00 | 6 of 15 benchmarks | 71.3% | 44.3% | 67.0% | 2.4% | 1513 | 1386 | |||||||||
| 164 | GPT-5.3-Codexopenai/gpt-5.3-codex | OpenAI | 54.4 | $1.75 / $14.00 | 1 of 15 benchmarks | 1409 | ||||||||||||||
| 165 | DeepSeek V4 Pro (xhigh)deepseek/deepseek-v4-pro:xhigh | DeepSeek | 54.4 | $1.168 / $2.336 | 2 of 15 benchmarks | 26.7% | 30.0 | |||||||||||||
| 166 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 54.3 | — | 1 of 15 benchmarks | 1452 | ||||||||||||||
| 167 | Step 3.5 Flashstepfun/step-3.5-flash | StepFun | 54.2 | $0.10 / $0.30 | 1 of 15 benchmarks | 1451 | ||||||||||||||
| 168 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 54.1 | $0.25 / $2.00 | 1 of 15 benchmarks | 64.7% | ||||||||||||||
| 169 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 53.9 | $1.168 / $2.336 | 3 of 15 benchmarks | 24.6% | 1502 | 1446 | ||||||||||||
| 170 | Claude Opus 4.1anthropic/claude-opus-4.1 | Anthropic | 53.9 | $15.00 / $75.00 | 4 of 15 benchmarks | 73.3% | 7.9% | 1505 | 1389 | |||||||||||
| 171 | DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepSeek | 53.6 | $0.25 / $0.95 | 1 of 15 benchmarks | 1448 | ||||||||||||||
| 172 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 53.5 | — | 2 of 15 benchmarks | 41.7% | 1385 | |||||||||||||
| 173 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 53.4 | $0.60 / $2.50 | 1 of 15 benchmarks | 63.4% | ||||||||||||||
| 174 | Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Qwen | 53.4 | $0.10 / $1.10 | 1 of 15 benchmarks | 1446 | ||||||||||||||
| 175 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 53.3 | $0.12 / $0.40 | 2 of 15 benchmarks | 1481 | 1362 | ||||||||||||||
| 176 | Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking | Qwen | 53.2 | $0.455 / $1.82 | 1 of 15 benchmarks | 1446 | ||||||||||||||
| 177 | GPT-5.5 (medium)openai/gpt-5.5:medium | OpenAI | 53.2 | $5.00 / $30.00 | 2 of 15 benchmarks | 54.0% | 36.6% | |||||||||||||
| 178 | R1deepseek/deepseek-r1 | DeepSeek | 53.1 | $0.70 / $2.50 | 1 of 15 benchmarks | 1445 | ||||||||||||||
| 179 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Meta | 52.9 | $0.35 / $1.50 | 2 of 15 benchmarks | 1481 | 1359 | |||||||||||||
| 180 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 52.9 | $0.5415 / $2.28 | 5 of 15 benchmarks | 76.7% | 22.2% | 2.4% | 1514 | 1509 | ||||||||||
| 181 | Grok 3 Betax-ai/grok-3-beta | xAI | 52.8 | — | 1 of 15 benchmarks | 1443 | ||||||||||||||
| 182 | Trinity Large Previewarcee-ai/trinity-large-preview | Arcee AI | 52.7 | — | 1 of 15 benchmarks | 1443 | ||||||||||||||
| 183 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 52.7 | $1.25 / $2.50 | 2 of 15 benchmarks | 1488 | 1354 | |||||||||||||
| 184 | o3 (medium)openai/o3:medium | OpenAI | 52.6 | $2.00 / $8.00 | 1 of 15 benchmarks | 62.3% | ||||||||||||||
| 185 | Gemini 3.7 Flash (low)google/gemini-3.7-flash:low | 52.6 | $0.375 / $1.875 | 2 of 15 benchmarks | 53.8% | 36.9% | ||||||||||||||
| 186 | Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Qwen | 52.5 | $0.23 / $2.30 | 1 of 15 benchmarks | 1442 | ||||||||||||||
| 187 | Qwen3 235B A22Bqwen/qwen3-235b-a22b | Qwen | 52.5 | $0.455 / $1.82 | 2 of 15 benchmarks | 65.9% | 1433 | |||||||||||||
| 188 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 52.3 | $0.0482 / $0.1931 | 1 of 15 benchmarks | 1440 | ||||||||||||||
| 189 | GPT-5.6 Terra (high)openai/gpt-5.6-terra:high | OpenAI | 52.2 | $1.00 / $6.00 | 2 of 15 benchmarks | 53.8% | 36.9% | |||||||||||||
| 190 | DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | DeepSeek | 52.1 | $0.27 / $0.95 | 1 of 15 benchmarks | 1439 | ||||||||||||||
| 191 | Hunyuan Vision 1.5 (thinking)tencent/hunyuan-vision-1.5:thinking | Tencent | 52.0 | — | 1 of 15 benchmarks | 1438 | ||||||||||||||
| 192 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 51.9 | $1.25 / $10.00 | 2 of 15 benchmarks | 66.0% | 1391 | |||||||||||||
| 193 | Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh | Anthropic | 51.9 | $5.00 / $25.00 | 1 of 15 benchmarks | 34.9% | ||||||||||||||
| 194 | Gemini 3.6 Flash (medium)google/gemini-3.6-flash:medium | 51.4 | $0.75 / $3.75 | 1 of 15 benchmarks | 34.4% | |||||||||||||||
| 195 | Grok 4 0709x-ai/grok-4-0709 | xAI | 51.3 | — | 1 of 15 benchmarks | 1435 | ||||||||||||||
| 196 | o3 Mini (low)openai/o3-mini:low | OpenAI | 51.2 | $1.10 / $4.40 | 1 of 15 benchmarks | 57.0% | ||||||||||||||
| 197 | DeepSeek V4 Flash 0423 (max)deepseek/deepseek-v4-flash:max | DeepSeek | 51.1 | $0.0643 / $0.1285 | 1 of 15 benchmarks | 53.3% | ||||||||||||||
| 198 | Muse Spark 1.1 (xhigh)meta/muse-spark-1.1:xhigh | Meta | 51.1 | $1.25 / $4.25 | 1 of 15 benchmarks | 53.3% | ||||||||||||||
| 199 | MiniMax M2.5minimax/minimax-m2.5 | MiniMax | 51.0 | $0.22 / $0.90 | 3 of 15 benchmarks | 68.3% | 1444 | 1384 | ||||||||||||
| 200 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 50.9 | $15.00 / $75.00 | 3 of 15 benchmarks | 70.7% | 46.9% | 1464 | ||||||||||||
| 201 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 50.9 | — | 1 of 15 benchmarks | 1433 | ||||||||||||||
| 202 | GPT-5.1openai/gpt-5.1 | OpenAI | 50.7 | $1.25 / $10.00 | 2 of 15 benchmarks | 1474 | 1341 | |||||||||||||
| 203 | Gemini 2.5 Progoogle/gemini-2.5-pro | 50.6 | $1.25 / $10.00 | 4 of 15 benchmarks | 57.6% | 73.6% | 1465 | 1226 | ||||||||||||
| 204 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 50.6 | $5.00 / $25.00 | 1 of 15 benchmarks | 12.7% | ||||||||||||||
| 205 | Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high | xAI | 50.6 | — | 2 of 15 benchmarks | 66.7% | 1390 | |||||||||||||
| 206 | Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | Moonshot AI | 50.4 | — | 2 of 15 benchmarks | 1486 | 1323 | |||||||||||||
| 207 | Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022 | Baidu | 50.3 | — | 1 of 15 benchmarks | 1432 | ||||||||||||||
| 208 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 50.1 | $0.269 / $0.40 | 3 of 15 benchmarks | 60.0% | 1475 | 1360 | ||||||||||||
| 209 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 50.1 | $0.75 / $4.50 | 3 of 15 benchmarks | 23.1% | 1497 | 1397 | ||||||||||||
| 210 | Kimi K2.5moonshotai/kimi-k2.5 | MoonshotAI | 50.1 | $0.57 / $2.85 | 3 of 15 benchmarks | 73.8% | 67.3% | 23.4% | ||||||||||||
| 211 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 50.1 | $0.25 / $2.00 | 1 of 15 benchmarks | 1431 | ||||||||||||||
| 212 | o1openai/o1 | OpenAI | 50.0 | $15.00 / $60.00 | 2 of 15 benchmarks | 64.6% | 1433 | |||||||||||||
| 213 | Claude Opus 4 (thinking)anthropic/claude-opus-4:thinking | Anthropic | 49.9 | $15.00 / $75.00 | 1 of 15 benchmarks | 56.6% | ||||||||||||||
| 214 | o4 Miniopenai/o4-mini | OpenAI | 49.8 | $1.10 / $4.40 | 3 of 15 benchmarks | 45.0% | 33.9% | 1433 | ||||||||||||
| 215 | DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | DeepSeek | 49.8 | $0.27 / $1.12 | 1 of 15 benchmarks | 1429 | ||||||||||||||
| 216 | Kimi K2.7 Code (none)moonshotai/kimi-k2.7-code:none | MoonshotAI | 49.7 | $0.71 / $3.50 | 1 of 15 benchmarks | 30.1% | ||||||||||||||
| 217 | GLM 4.5 Airz-ai/glm-4.5-air | Z.ai | 49.6 | $0.13 / $0.85 | 1 of 15 benchmarks | 1426 | ||||||||||||||
| 218 | Devstral Small 2512mistralai/devstral-small-2512 | Mistral AI | 49.6 | — | 1 of 15 benchmarks | 56.4% | ||||||||||||||
| 219 | GLM 4.7 Flashz-ai/glm-4.7-flash | Z.ai | 49.5 | $0.06 / $0.40 | 1 of 15 benchmarks | 1424 | ||||||||||||||
| 220 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 49.4 | $0.29 / $2.40 | 2 of 15 benchmarks | 1459 | 1358 | |||||||||||||
| 221 | Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview | Tencent | 49.4 | — | 2 of 15 benchmarks | 1461 | 1356 | |||||||||||||
| 222 | Solar Pro 4upstage/solar-pro4 | Upstage | 49.3 | $0.03 / $0.12 | 2 of 15 benchmarks | 1450 | 1371 | |||||||||||||
| 223 | GPT-5.5 (low)openai/gpt-5.5:low | OpenAI | 49.3 | $5.00 / $30.00 | 4 of 15 benchmarks | 27.0% | 30.6% | 32.6% | 38.0 | |||||||||||
| 224 | Nemotron 3.5 Lightning 30B A3Bnvidia/nemotron-3.5-lightning-30b-a3b | NVIDIA | 49.2 | — | 1 of 15 benchmarks | 1422 | ||||||||||||||
| 225 | MiniMax M2.1minimax/minimax-m2.1 | MiniMax | 49.2 | $0.30 / $1.20 | 2 of 15 benchmarks | 1440 | 1387 | |||||||||||||
| 226 | Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Qwen | 49.1 | $0.15 / $1.20 | 1 of 15 benchmarks | 1421 | ||||||||||||||
| 227 | GLM 4.6Vz-ai/glm-4.6v | Z.ai | 49.0 | $0.30 / $0.90 | 1 of 15 benchmarks | 1417 | ||||||||||||||
| 228 | Grok 4.1 (thinking)x-ai/grok-4.1:thinking | xAI | 48.9 | — | 2 of 15 benchmarks | 1499 | 1210 | |||||||||||||
| 229 | GLM 4.5z-ai/glm-4.5 | Z.ai | 48.8 | $0.60 / $2.20 | 2 of 15 benchmarks | 54.2% | 1455 | |||||||||||||
| 230 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 48.8 | — | 6 of 15 benchmarks | 62.8% | 51.3% | 24.9% | 25.3% | 36.4% | 1435 | |||||||||
| 231 | MiniMax M1minimax/minimax-m1 | MiniMax | 48.7 | $0.55 / $2.20 | 1 of 15 benchmarks | 1416 | ||||||||||||||
| 232 | Claude Sonnet 5anthropic/claude-sonnet-5 | Anthropic | 48.6 | $2.00 / $10.00 | 2 of 15 benchmarks | 25.6% | 27.0 | |||||||||||||
| 233 | Claude Sonnet 4 (thinking)anthropic/claude-sonnet-4:thinking | Anthropic | 48.5 | $3.00 / $15.00 | 1 of 15 benchmarks | 56.0% | ||||||||||||||
| 234 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 48.4 | — | 1 of 15 benchmarks | 1412 | ||||||||||||||
| 235 | Claude Opus 4.7 (low)anthropic/claude-opus-4.7:low | Anthropic | 48.4 | $5.00 / $25.00 | 1 of 15 benchmarks | 27.6% | ||||||||||||||
| 236 | Ling Flash 2.0inclusionai/ling-flash-2.0 | inclusionAI | 48.3 | — | 1 of 15 benchmarks | 1411 | ||||||||||||||
| 237 | Composer 2.5cursor/composer-2.5 | Cursor | 48.3 | — | 1 of 15 benchmarks | 33.8% | ||||||||||||||
| 238 | Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | 48.1 | $0.325 / $1.95 | 4 of 15 benchmarks | 57.9% | 19.9% | 1495 | 1460 | |||||||||||
| 239 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 48.1 | $1.50 / $7.50 | 2 of 15 benchmarks | 1479 | 1265 | |||||||||||||
| 240 | INTELLECT-3prime-intellect/intellect-3 | Prime Intellect | 48.1 | — | 1 of 15 benchmarks | 1409 | ||||||||||||||
| 241 | Step 3stepfun/step-3 | StepFun | 48.0 | — | 1 of 15 benchmarks | 1408 | ||||||||||||||
| 242 | Inklingthinkingmachines/inkling | Thinking Machines | 48.0 | $0.95 / $4.05 | 3 of 15 benchmarks | 14.0% | 1494 | 1405 | ||||||||||||
| 243 | GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh | OpenAI | 47.9 | $0.75 / $4.50 | 1 of 15 benchmarks | 27.0% | ||||||||||||||
| 244 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 47.9 | — | 3 of 15 benchmarks | 69.6% | 1457 | 1273 | ||||||||||||
| 245 | Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | NVIDIA | 47.9 | $0.085 / $0.40 | 1 of 15 benchmarks | 1408 | ||||||||||||||
| 246 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 47.9 | $0.195 / $1.56 | 2 of 15 benchmarks | 1450 | 1358 | |||||||||||||
| 247 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 47.8 | $0.71 / $3.50 | 2 of 15 benchmarks | 30.5% | 1473 | |||||||||||||
| 248 | Qwen3 32Bqwen/qwen3-32b | Qwen | 47.7 | $0.08 / $0.28 | 1 of 15 benchmarks | 1407 | ||||||||||||||
| 249 | Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | Qwen | 47.7 | $0.07 / $0.28 | 1 of 15 benchmarks | 51.6% | ||||||||||||||
| 250 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 47.6 | $0.60 / $1.80 | 1 of 15 benchmarks | 1405 | ||||||||||||||
| 251 | Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | NVIDIA | 47.5 | — | 1 of 15 benchmarks | 1404 | ||||||||||||||
| 252 | Qwen 2.5 (max)qwen/qwen-2.5:max | Qwen | 47.3 | — | 1 of 15 benchmarks | 1403 | ||||||||||||||
| 253 | Hunyuan T1tencent/hunyuan-t1 | Tencent | 47.2 | — | 1 of 15 benchmarks | 1399 | ||||||||||||||
| 254 | GPT-4.1openai/gpt-4.1 | OpenAI | 47.2 | $2.00 / $8.00 | 3 of 15 benchmarks | 48.5% | 31.1% | 1456 | ||||||||||||
| 255 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 47.0 | $0.10 / $0.40 | 1 of 15 benchmarks | 1397 | |||||||||||||||
| 256 | GPT-5.6 Sol (low)openai/gpt-5.6-sol:low | OpenAI | 47.0 | $5.00 / $30.00 | 2 of 15 benchmarks | 45.4% | 35.4% | |||||||||||||
| 257 | Devstral Small 2505mistralai/devstral-small-2505 | Mistral AI | 46.9 | — | 1 of 15 benchmarks | 46.8% | ||||||||||||||
| 258 | Nova 2 Liteamazon/nova-2-lite-v1 | Amazon | 46.9 | $0.30 / $2.50 | 1 of 15 benchmarks | 1395 | ||||||||||||||
| 259 | Laguna M.1poolside/laguna-m.1 | Poolside | 46.9 | — | 1 of 15 benchmarks | 1347 | ||||||||||||||
| 260 | Hunyuan TurboStencent/hunyuan-turbos | Tencent | 46.8 | — | 1 of 15 benchmarks | 1394 | ||||||||||||||
| 261 | DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | DeepSeek | 46.7 | $0.27 / $0.41 | 2 of 15 benchmarks | 1465 | 1272 | |||||||||||||
| 262 | Composer 2.5 (none)cursor/composer-2.5:none | Cursor | 46.6 | — | 1 of 15 benchmarks | 25.6% | ||||||||||||||
| 263 | Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1 | NVIDIA | 46.5 | — | 1 of 15 benchmarks | 1391 | ||||||||||||||
| 264 | Ring Flash 2.0inclusionai/ring-flash-2.0 | inclusionAI | 46.4 | — | 1 of 15 benchmarks | 1390 | ||||||||||||||
| 265 | GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | 46.3 | $1.75 / $14.00 | 3 of 15 benchmarks | 72.8% | 66.3% | 1338 | ||||||||||||
| 266 | Kimi K2 Instructmoonshotai/kimi-k2-instruct | Moonshot AI | 46.2 | — | 1 of 15 benchmarks | 43.8% | ||||||||||||||
| 267 | GLM 5.2 (none)z-ai/glm-5.2:none | Z.ai | 46.2 | $0.462 / $1.452 | 1 of 15 benchmarks | 24.5% | ||||||||||||||
| 268 | Command A (03-2025)cohere/command-a-03-2025 | Cohere | 46.0 | — | 1 of 15 benchmarks | 1390 | ||||||||||||||
| 269 | o3 Miniopenai/o3-mini | OpenAI | 45.8 | $1.10 / $4.40 | 4 of 15 benchmarks | 42.4% | 32.3% | 63.0% | 1416 | |||||||||||
| 270 | Mimo v2 Flashxiaomi/mimo-v2-flash | Xiaomi | 45.8 | — | 2 of 15 benchmarks | 1446 | 1330 | |||||||||||||
| 271 | Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max | Anthropic | 45.8 | $3.00 / $15.00 | 1 of 15 benchmarks | 24.3% | ||||||||||||||
| 272 | Magistral Medium 2506mistralai/magistral-medium-2506 | Mistral AI | 45.7 | — | 1 of 15 benchmarks | 1387 | ||||||||||||||
| 273 | GPT-5.1-Codexopenai/gpt-5.1-codex | OpenAI | 45.7 | $1.25 / $10.00 | 1 of 15 benchmarks | 1336 | ||||||||||||||
| 274 | Nova Premier 1.0amazon/nova-premier-v1 | Amazon | 45.6 | $2.50 / $12.50 | 1 of 15 benchmarks | 42.4% | ||||||||||||||
| 275 | O1 Miniopenai/o1-mini | OpenAI | 45.5 | — | 1 of 15 benchmarks | 1387 | ||||||||||||||
| 276 | Grok 3 Mini Betax-ai/grok-3-mini-beta | xAI | 45.4 | — | 1 of 15 benchmarks | 1386 | ||||||||||||||
| 277 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Qwen | 45.3 | $0.12 / $0.50 | 1 of 15 benchmarks | 1386 | ||||||||||||||
| 278 | Grok 4.6 (low)x-ai/grok-4.6:low | SpaceXAI | 45.2 | $2.00 / $6.00 | 1 of 15 benchmarks | 41.6% | ||||||||||||||
| 279 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 45.0 | — | 2 of 15 benchmarks | 1468 | 1230 | |||||||||||||
| 280 | GLM 4.6z-ai/glm-4.6 | Z.ai | 44.9 | $0.55 / $2.20 | 3 of 15 benchmarks | 55.4% | 1458 | 1340 | ||||||||||||
| 281 | Claude Opus 4.6 (medium)anthropic/claude-opus-4.6:medium | Anthropic | 44.9 | $5.00 / $25.00 | 1 of 15 benchmarks | 23.7% | ||||||||||||||
| 282 | QwQ 32Bqwen/qwq-32b | Qwen | 44.9 | — | 1 of 15 benchmarks | 1384 | ||||||||||||||
| 283 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 44.8 | — | 2 of 15 benchmarks | 1461 | 1240 | |||||||||||||
| 284 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 44.7 | $0.05 / $0.40 | 1 of 15 benchmarks | 1384 | ||||||||||||||
| 285 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 44.6 | $0.10 / $0.40 | 1 of 15 benchmarks | 1384 | |||||||||||||||
| 286 | Claude Sonnet 5 (medium)anthropic/claude-sonnet-5:medium | Anthropic | 44.6 | $2.00 / $10.00 | 2 of 15 benchmarks | 39.8% | 35.2% | |||||||||||||
| 287 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 44.5 | $0.25 / $1.50 | 2 of 15 benchmarks | 1457 | 1254 | ||||||||||||||
| 288 | OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct | Allen Institute for AI | 44.4 | — | 1 of 15 benchmarks | 1382 | ||||||||||||||
| 289 | Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low | Anthropic | 44.3 | $5.00 / $25.00 | 2 of 15 benchmarks | 40.8% | 35.1% | |||||||||||||
| 290 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 44.3 | $0.10 / $0.40 | 1 of 15 benchmarks | 1374 | ||||||||||||||
| 291 | Laguna XS.2poolside/laguna-xs.2 | Poolside | 44.2 | — | 1 of 15 benchmarks | 1303 | ||||||||||||||
| 292 | gpt-oss-20bopenai/gpt-oss-20b | OpenAI | 44.0 | $0.03 / $0.13 | 1 of 15 benchmarks | 1370 | ||||||||||||||
| 293 | Devstral Small 2507mistralai/devstral-small-2507 | Mistral AI | 43.9 | — | 1 of 15 benchmarks | 38.0% | ||||||||||||||
| 294 | DeepSeek V4 Flash 0423 (xhigh)deepseek/deepseek-v4-flash:xhigh | DeepSeek | 43.8 | $0.0643 / $0.1285 | 2 of 15 benchmarks | 20.9% | 27.0 | |||||||||||||
| 295 | Mercuryinception/mercury | Inception | 43.8 | — | 1 of 15 benchmarks | 1367 | ||||||||||||||
| 296 | Gemini 3.6 Flash (low)google/gemini-3.6-flash:low | 43.6 | $0.75 / $3.75 | 1 of 15 benchmarks | 22.8% | |||||||||||||||
| 297 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 43.5 | $0.05 / $0.40 | 1 of 15 benchmarks | 34.8% | ||||||||||||||
| 298 | OLMo 3 32B Thinkallenai/olmo-3-32b-think | Allen Institute for AI | 43.5 | — | 1 of 15 benchmarks | 1364 | ||||||||||||||
| 299 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 43.4 | $1.00 / $5.00 | 3 of 15 benchmarks | 64.7% | 1479 | 1326 | ||||||||||||
| 300 | Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | NVIDIA | 43.3 | — | 1 of 15 benchmarks | 1363 | ||||||||||||||
| 301 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | NVIDIA | 43.2 | $0.05 / $0.20 | 1 of 15 benchmarks | 1362 | ||||||||||||||
| 302 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 43.2 | $2.50 / $10.00 | 1 of 15 benchmarks | 31.0% | ||||||||||||||
| 303 | Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium | Anthropic | 43.1 | $3.00 / $15.00 | 1 of 15 benchmarks | 21.1% | ||||||||||||||
| 304 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 42.9 | — | 1 of 15 benchmarks | 1362 | ||||||||||||||
| 305 | GPT-5.4 Mini (medium)openai/gpt-5.4-mini:medium | OpenAI | 42.7 | $0.75 / $4.50 | 1 of 15 benchmarks | 20.8% | ||||||||||||||
| 306 | Gemma 3 27Bgoogle/gemma-3-27b-it | 42.7 | $0.08 / $0.45 | 1 of 15 benchmarks | 1358 | |||||||||||||||
| 307 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 42.5 | — | 1 of 15 benchmarks | 1356 | |||||||||||||||
| 308 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 42.4 | — | 1 of 15 benchmarks | 1356 | ||||||||||||||
| 309 | KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1 | Kwaipilot | 42.4 | — | 1 of 15 benchmarks | 1255 | ||||||||||||||
| 310 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Qwen | 42.3 | $0.36 / $0.40 | 1 of 15 benchmarks | 1356 | ||||||||||||||
| 311 | DeepSeek V4 Flash 0731 (high)deepseek/deepseek-v4-flash-0731:high | DeepSeek | 42.2 | $0.14 / $0.28 | 1 of 15 benchmarks | 18.8% | ||||||||||||||
| 312 | Mistral Large 2407mistralai/mistral-large-2407 | Mistral | 42.0 | $2.00 / $6.00 | 1 of 15 benchmarks | 1354 | ||||||||||||||
| 313 | Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking | Xiaomi | 41.9 | — | 2 of 15 benchmarks | 1431 | 1292 | |||||||||||||
| 314 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 41.8 | $5.00 / $15.00 | 4 of 15 benchmarks | 38.8% | 31.3% | 12.0% | 1369 | |||||||||||
| 315 | Claude Opus 4.6 (low)anthropic/claude-opus-4.6:low | Anthropic | 41.8 | $5.00 / $25.00 | 1 of 15 benchmarks | 18.3% | ||||||||||||||
| 316 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 41.7 | — | 1 of 15 benchmarks | 1352 | ||||||||||||||
| 317 | GPT-5.6 Terra (medium)openai/gpt-5.6-terra:medium | OpenAI | 41.6 | $1.00 / $6.00 | 2 of 15 benchmarks | 35.1% | 33.9% | |||||||||||||
| 318 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Qwen | 41.5 | $0.225 / $1.80 | 2 of 15 benchmarks | 1435 | 1250 | |||||||||||||
| 319 | GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | OpenAI | 41.5 | $0.25 / $2.00 | 1 of 15 benchmarks | 1244 | ||||||||||||||
| 320 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 41.4 | $0.30 / $2.50 | 3 of 15 benchmarks | 28.7% | 61.9% | 1424 | |||||||||||||
| 321 | DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none | DeepSeek | 41.4 | $1.168 / $2.336 | 1 of 15 benchmarks | 17.6% | ||||||||||||||
| 322 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 41.3 | — | 1 of 15 benchmarks | 1347 | |||||||||||||||
| 323 | Mistral Large 2411mistralai/mistral-large-2411 | Mistral AI | 41.1 | — | 1 of 15 benchmarks | 1346 | ||||||||||||||
| 324 | DeepSeek V3deepseek/deepseek-chat | DeepSeek | 41.1 | $0.2574 / $1.0287 | 3 of 15 benchmarks | 36.7% | 27.2% | 1388 | ||||||||||||
| 325 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 41.1 | $0.40 / $1.60 | 2 of 15 benchmarks | 23.9% | 1433 | |||||||||||||
| 326 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Meta | 41.0 | $0.10 / $0.32 | 1 of 15 benchmarks | 1346 | ||||||||||||||
| 327 | Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Qwen | 40.9 | $0.065 / $0.26 | 2 of 15 benchmarks | 1437 | 1238 | |||||||||||||
| 328 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 40.9 | — | 1 of 15 benchmarks | 1343 | ||||||||||||||
| 329 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 40.7 | — | 1 of 15 benchmarks | 1343 | |||||||||||||||
| 330 | Qwen 2.5qwen/qwen-2.5 | Qwen | 40.6 | — | 2 of 15 benchmarks | 40.2% | 24.7% | |||||||||||||
| 331 | DeepSeek V3.2deepseek/deepseek-v3.2 | DeepSeek | 40.6 | $0.269 / $0.40 | 3 of 15 benchmarks | 59.0% | 1470 | 1325 | ||||||||||||
| 332 | Claude Opus 4.8 (medium)anthropic/claude-opus-4.8:medium | Anthropic | 40.6 | $5.00 / $25.00 | 4 of 15 benchmarks | 48.7% | 40.5% | 20.9% | 19.0 | |||||||||||
| 333 | Claude Sonnet 4.6 (low)anthropic/claude-sonnet-4.6:low | Anthropic | 40.5 | $3.00 / $15.00 | 1 of 15 benchmarks | 15.0% | ||||||||||||||
| 334 | OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think | Allen Institute for AI | 40.3 | — | 1 of 15 benchmarks | 1338 | ||||||||||||||
| 335 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Meta | 40.2 | $0.40 / $0.40 | 1 of 15 benchmarks | 1333 | ||||||||||||||
| 336 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 40.1 | — | 1 of 15 benchmarks | 1318 | ||||||||||||||
| 337 | MiniMax M3 (none)minimax/minimax-m3:none | MiniMax | 40.1 | $0.30 / $1.20 | 1 of 15 benchmarks | 14.7% | ||||||||||||||
| 338 | Gemma 3 12Bgoogle/gemma-3-12b-it | 39.9 | $0.05 / $0.15 | 1 of 15 benchmarks | 1316 | |||||||||||||||
| 339 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 39.8 | — | 1 of 15 benchmarks | 1316 | |||||||||||||||
| 340 | Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501 | Mistral AI | 39.6 | — | 1 of 15 benchmarks | 1312 | ||||||||||||||
| 341 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 39.5 | — | 1 of 15 benchmarks | 1309 | |||||||||||||||
| 342 | Gemma 3n E4Bgoogle/gemma-3n-e4b-it | 39.4 | — | 1 of 15 benchmarks | 1308 | |||||||||||||||
| 343 | MiniMax M2minimax/minimax-m2 | MiniMax | 39.2 | $0.255 / $1.02 | 3 of 15 benchmarks | 61.0% | 1385 | 1297 | ||||||||||||
| 344 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 39.2 | — | 1 of 15 benchmarks | 1306 | ||||||||||||||
| 345 | Qwen3.7 Plus (none)qwen/qwen3.7-plus:none | Qwen | 39.2 | $0.32 / $1.28 | 1 of 15 benchmarks | 10.2% | ||||||||||||||
| 346 | Claude Sonnet 5 (low)anthropic/claude-sonnet-5:low | Anthropic | 39.2 | $2.00 / $10.00 | 2 of 15 benchmarks | 30.5% | 28.7% | |||||||||||||
| 347 | GPT 4 1106 Previewopenai/gpt-4-1106-preview | OpenAI | 39.1 | — | 4 of 15 benchmarks | 22.4% | 28.3% | 12.5% | 1340 | |||||||||||
| 348 | Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning | xAI | 39.0 | — | 2 of 15 benchmarks | 1436 | 1161 | |||||||||||||
| 349 | Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0 | Amazon | 39.0 | — | 1 of 15 benchmarks | 1289 | ||||||||||||||
| 350 | Command R (08-2024)cohere/command-r-08-2024 | Cohere | 38.8 | $0.15 / $0.60 | 1 of 15 benchmarks | 1281 | ||||||||||||||
| 351 | SWE-1.6 (none)cognition/swe-1.6:none | Cognition | 38.7 | — | 1 of 15 benchmarks | 9.4% | ||||||||||||||
| 352 | Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | 38.7 | $2.50 / $10.00 | 1 of 15 benchmarks | 1280 | ||||||||||||||
| 353 | Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | 38.7 | $0.22 / $0.85 | 2 of 15 benchmarks | 1414 | 1239 | |||||||||||||
| 354 | OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct | Allen Institute for AI | 38.5 | — | 1 of 15 benchmarks | 1280 | ||||||||||||||
| 355 | Grok Code Fast 1x-ai/grok-code-fast-1 | xAI | 38.5 | — | 1 of 15 benchmarks | 1164 | ||||||||||||||
| 356 | Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | 38.4 | $2.00 / $6.00 | 1 of 15 benchmarks | 1277 | ||||||||||||||
| 357 | GPT-5.4 Mini (low)openai/gpt-5.4-mini:low | OpenAI | 38.3 | $0.75 / $4.50 | 1 of 15 benchmarks | 9.3% | ||||||||||||||
| 358 | Gemma 3 4Bgoogle/gemma-3-4b-it | 38.3 | — | 1 of 15 benchmarks | 1274 | |||||||||||||||
| 359 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 38.1 | — | 1 of 15 benchmarks | 1272 | |||||||||||||||
| 360 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Meta | 38.0 | $0.05 / $0.08 | 1 of 15 benchmarks | 1260 | ||||||||||||||
| 361 | Devstral Medium 2507mistralai/devstral-medium-2507 | Mistral AI | 37.9 | — | 1 of 15 benchmarks | 1080 | ||||||||||||||
| 362 | Mistral Medium 3.5 (none)mistralai/mistral-medium-3-5:none | Mistral | 37.9 | $1.50 / $7.50 | 1 of 15 benchmarks | 8.0% | ||||||||||||||
| 363 | GPT-4oopenai/gpt-4o | OpenAI | 37.9 | $2.50 / $10.00 | 1 of 15 benchmarks | 12.2% | ||||||||||||||
| 364 | QwQ 32B Previewqwen/qwq-32b-preview | Qwen | 37.9 | — | 1 of 15 benchmarks | 1173 | ||||||||||||||
| 365 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 37.8 | $0.03 / $0.17 | 2 of 15 benchmarks | 26.0% | 1390 | |||||||||||||
| 366 | Devstral 2mistralai/devstral-2 | Mistral AI | 37.2 | — | 2 of 15 benchmarks | 53.8% | 1194 | |||||||||||||
| 367 | GPT-5 Miniopenai/gpt-5-mini | OpenAI | 36.9 | $0.25 / $2.00 | 2 of 15 benchmarks | 56.2% | 39.7% | |||||||||||||
| 368 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 36.9 | $2.50 / $10.00 | 5 of 15 benchmarks | 27.0% | 39.7% | 30.4% | 29.5% | 1360 | ||||||||||
| 369 | GPT-5.6 Luna (medium)openai/gpt-5.6-luna:medium | OpenAI | 35.7 | $0.10 / $0.60 | 2 of 15 benchmarks | 11.3% | 25.7% | |||||||||||||
| 370 | Mercury 2inception/mercury-2 | Inception | 35.6 | $0.25 / $0.75 | 2 of 15 benchmarks | 1394 | 1166 | |||||||||||||
| 371 | Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high | Anthropic | 35.4 | $3.00 / $15.00 | 2 of 15 benchmarks | 29.9% | 23.5% | |||||||||||||
| 372 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 35.4 | — | 2 of 15 benchmarks | 21.0% | 1373 | |||||||||||||
| 373 | GPT-5.6 Terra (low)openai/gpt-5.6-terra:low | OpenAI | 35.3 | $1.00 / $6.00 | 2 of 15 benchmarks | 24.1% | 24.1% | |||||||||||||
| 374 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 35.0 | — | 4 of 15 benchmarks | 15.8% | 26.3% | 10.5% | 1355 | |||||||||||
| 375 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 34.4 | — | 2 of 15 benchmarks | 13.5% | 1365 | ||||||||||||||
| 376 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 34.1 | $10.00 / $30.00 | 2 of 15 benchmarks | 28.7% | 1347 | |||||||||||||
| 377 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 33.7 | — | 2 of 15 benchmarks | 9.1% | 1362 | |||||||||||||
| 378 | Claude 2anthropic/claude-2 | Anthropic | 33.7 | — | 3 of 15 benchmarks | 4.4% | 3.0% | 2.0% | ||||||||||||
| 379 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 33.2 | $0.15 / $0.60 | 2 of 15 benchmarks | 27.5% | 1349 | |||||||||||||
| 380 | Granite 4.1 8Bibm-granite/granite-4.1-8b | IBM | 32.3 | $0.05 / $0.10 | 2 of 15 benchmarks | 1353 | 1192 | |||||||||||||
| 381 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 31.6 | — | 2 of 15 benchmarks | 9.0% | 1342 | |||||||||||||
| 382 | GPT-5.6 Luna (low)openai/gpt-5.6-luna:low | OpenAI | 30.7 | $0.10 / $0.60 | 2 of 15 benchmarks | 1.6% | 15.4% | |||||||||||||
| 383 | Claude Opus 4.7 (medium)anthropic/claude-opus-4.7:medium | Anthropic | 29.6 | $5.00 / $25.00 | 3 of 15 benchmarks | 28.9% | 7.0% | 7.0 | ||||||||||||
| 384 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 29.2 | $2.00 / $12.00 | 2 of 15 benchmarks | 11.7% | 8.9% | ||||||||||||||
| 385 | GLM 5.2 (high)z-ai/glm-5.2:high | Z.ai | 29.2 | $0.462 / $1.452 | 3 of 15 benchmarks | 36.3% | 17.4% | 18.0 | ||||||||||||
| 386 | SWE Llamaprinceton-nlp/swe-llama | Princeton NLP | 29.0 | — | 3 of 15 benchmarks | 1.4% | 1.3% | 0.7% | ||||||||||||
| 387 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 27.8 | $0.25 / $1.25 | 3 of 15 benchmarks | 40.6% | 20.2% | 1301 | ||||||||||||
| 388 | SWE Llama 13Bprinceton-nlp/swe-llama-13b | Princeton NLP | 27.4 | — | 3 of 15 benchmarks | 1.2% | 1.0% | 0.7% | ||||||||||||
| 389 | GPT 3.5openai/gpt-3.5 | OpenAI | 22.7 | — | 3 of 15 benchmarks | 0.4% | 0.3% | 0.2% |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 15 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- DeepSWE v1.1
Pass@1 across all 113 DeepSWE v1.1 tasks, run by Datacurve with mini-swe-agent through Pier; reasoning-effort configurations remain separate.
- Epoch AI Benchmarking HubCC BY 4.0
Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- FrontierCode 1.1
Main weighted rubric Scores on FrontierCode 1.1's 100 hardest tasks, run by Cognition across five trials; results identify model, harness, and effort.
- FrontierSWE
Mean@5 Dominance across FrontierSWE's complete 17-task cohort; results identify the evaluated model and harness.
- LiveCodeBenchMIT
Maintainer-published Code Generation pass@1 over the 454-problem release_v6 window (2024-08-01 through 2025-05-01), from the stable 2025-08-01 snapshot; it does not cover newer frontier families.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.
- SWE-bench
Resolve rates published by the SWE-bench maintainers.
- Warden
Security review results published by Warden.