---
title: "LLM leaderboard"
description: "Every LLM benchmark this pipeline can legally republish, normalised to percentiles and ranked per category, with coverage shown on every row."
url: "https://aldena.ai/llm-leaderboard"
---

# LLM leaderboard

The overall table is a composite of four text categories: coding, agentic tool use, math and vision. Media categories are reachable by tab and stay out of the overall number, because ranking a speech model against a coding model produces a figure that means nothing.

Coverage is printed on every row. A model the benchmarks have only partly reached still ranks, on the results it does have, with a mark on its row saying so, and its score is pulled toward the mean of the broadly benchmarked models rather than padded with blanks.

Aldena runs these models inside your team rooms. [See what each one costs](https://aldena.ai/models).

| rank | model | vendor | score | coverage (categories) | price (in / out) | Agentic tool use (composite) | Coding (composite) | Math (composite) | Vision (composite) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 (max) | Anthropic | 68.6 | 3 of 4 | $5.00 / $25.00 | 75.9 | 86.1 | 80.5 | — |
| 2 | Claude Opus 5 (high) (partial coverage) | Anthropic | 67.9 | 4 of 4 | $5.00 / $25.00 | 78.0 | 82.2 | 65.8 | 81.0 |
| 3 | Claude Fable 5 (partial coverage) | Anthropic | 65.9 | 4 of 4 | $10.00 / $50.00 | 63.3 | 81.8 | 65.9 | 84.0 |
| 4 | Kimi K3 (max) | MoonshotAI | 65.8 | 3 of 4 | $3.00 / $15.00 | 75.7 | 79.1 | 74.1 | — |
| 5 | Claude Fable 5 (max) | Anthropic | 65.2 | 2 of 4 | $10.00 / $50.00 | — | 78.7 | 82.0 | — |
| 6 | Claude Opus 4.6 | Anthropic | 65.2 | 4 of 4 | $5.00 / $25.00 | 74.0 | 75.8 | 65.3 | 75.9 |
| 7 | Claude Fable 5 (high) (partial coverage) | Anthropic | 65.2 | 3 of 4 | $10.00 / $50.00 | 81.0 | 78.8 | 65.8 | — |
| 8 | GPT-5.6 Sol (xhigh) (partial coverage) | OpenAI | 65.2 | 4 of 4 | $5.00 / $30.00 | 77.1 | 80.4 | 62.5 | 70.7 |
| 9 | Claude Opus 4.7 (partial coverage) | Anthropic | 64.3 | 4 of 4 | $5.00 / $25.00 | 68.8 | 71.5 | 64.1 | 80.9 |
| 10 | Grok 4.5 (partial coverage) | SpaceXAI | 64.2 | 4 of 4 | $2.00 / $6.00 | 69.4 | 74.6 | 62.8 | 78.3 |
| 11 | GPT-5.6 Sol (max) | OpenAI | 64.1 | 2 of 4 | $5.00 / $30.00 | — | 77.0 | 79.3 | — |
| 12 | Qwen3.8 Max (partial coverage) | Qwen | 64.1 | 4 of 4 | $2.00 / $6.00 | 60.2 | 76.7 | 65.5 | 81.7 |
| 13 | GPT-5.5 (high) (partial coverage) | OpenAI | 64.0 | 4 of 4 | $5.00 / $30.00 | 73.3 | 69.2 | 64.2 | 76.9 |
| 14 | GPT-5.4 (high) | OpenAI | 63.3 | 4 of 4 | $2.50 / $15.00 | 64.3 | 66.4 | 78.8 | 70.1 |
| 15 | GPT-5.5 (xhigh) (partial coverage) | OpenAI | 63.0 | 3 of 4 | $5.00 / $30.00 | 73.3 | 70.5 | 70.9 | — |
| 16 | GPT-5.5 (partial coverage) | OpenAI | 62.6 | 4 of 4 | $5.00 / $30.00 | 68.8 | 64.3 | 64.8 | 77.6 |
| 17 | Gemini 3 Pro (partial coverage) | Google | 61.6 | 3 of 4 | — | — | 65.0 | 62.6 | 79.9 |
| 18 | Muse Spark (partial coverage) | Meta | 61.4 | 4 of 4 | — | 60.5 | 69.1 | 64.4 | 74.4 |
| 19 | GPT 5.5 Pre Release (xhigh) (partial coverage) | OpenAI | 61.2 | 2 of 4 | — | — | 70.8 | 73.7 | — |
| 20 | Claude Opus 4.7 (high) (partial coverage) | Anthropic | 61.1 | 3 of 4 | $5.00 / $25.00 | 69.1 | 71.0 | 64.9 | — |
| 21 | Muse Spark 1.2 (xhigh) | Meta | 60.9 | 2 of 4 | $1.25 / $4.25 | — | 68.9 | — | 74.4 |
| 22 | Claude Opus 4.8 (max) (partial coverage) | Anthropic | 60.9 | 3 of 4 | $5.00 / $25.00 | 60.5 | 68.9 | 74.8 | — |
| 23 | Muse Spark 1.1 (partial coverage) | Meta | 60.8 | 4 of 4 | $1.25 / $4.25 | 60.2 | 73.5 | 63.6 | 67.3 |
| 24 | Claude Fable 5 (xhigh) (partial coverage) | Anthropic | 60.8 | 2 of 4 | $10.00 / $50.00 | 66.8 | 76.3 | — | — |
| 25 | Claude Opus 4.6 (thinking) | Anthropic | 60.5 | 1 of 4 | $5.00 / $25.00 | — | — | — | 81.3 |
| 26 | Claude Opus 4.8 (partial coverage) | Anthropic | 60.4 | 4 of 4 | $5.00 / $25.00 | 57.2 | 72.0 | 61.4 | 71.6 |
| 27 | Claude Opus 4.7 (thinking) | Anthropic | 60.3 | 1 of 4 | $5.00 / $25.00 | — | — | — | 80.5 |
| 28 | GPT-5.6 Terra (xhigh) (partial coverage) | OpenAI | 60.0 | 4 of 4 | $1.00 / $6.00 | 60.2 | 66.2 | 63.4 | 69.8 |
| 29 | Grok 4.6 (high) (partial coverage) | SpaceXAI | 59.9 | 2 of 4 | $2.00 / $6.00 | — | 75.5 | 63.8 | — |
| 30 | Claude Opus 5 (xhigh) (partial coverage) | Anthropic | 59.7 | 2 of 4 | $5.00 / $25.00 | 65.6 | 73.0 | — | — |
| 31 | GPT-5.6 Terra (max) (partial coverage) | OpenAI | 59.5 | 3 of 4 | $1.00 / $6.00 | 52.9 | 67.2 | 77.0 | — |
| 32 | Gemini 3.7 Flash (high) | Google | 59.4 | 3 of 4 | $0.375 / $1.875 | 57.8 | 69.8 | 69.2 | — |
| 33 | Claude Opus 4.8 (thinking) | Anthropic | 59.4 | 1 of 4 | $5.00 / $25.00 | — | — | — | 77.9 |
| 34 | DeepSeek V4 Pro (high) (partial coverage) | DeepSeek | 59.0 | 2 of 4 | $1.168 / $2.336 | — | 68.8 | 67.0 | — |
| 35 | Gemini 3.1 Pro Preview | Google | 59.0 | 4 of 4 | $2.00 / $12.00 | 44.4 | 58.8 | 72.6 | 77.9 |
| 36 | Claude Sonnet 5 (high) (partial coverage) | Anthropic | 58.9 | 4 of 4 | $2.00 / $10.00 | 57.1 | 63.5 | 60.4 | 72.0 |
| 37 | Gemini 3 Flash Preview (high) (partial coverage) | Google | 58.6 | 3 of 4 | $0.50 / $3.00 | 65.7 | 65.7 | 61.6 | — |
| 38 | Qwen3.6 Max Preview | Qwen | 58.6 | 2 of 4 | $1.027 / $6.162 | — | 70.0 | 64.1 | — |
| 39 | GPT-5.2 Chat (partial coverage) | OpenAI | 58.6 | 3 of 4 | $1.75 / $14.00 | — | 67.1 | 58.1 | 67.3 |
| 40 | Claude Opus 4.5 (medium) (partial coverage) | Anthropic | 58.5 | 2 of 4 | $5.00 / $25.00 | 63.8 | 70.0 | — | — |
| 41 | Gemini 3.6 Flash | Google | 58.5 | 1 of 4 | $0.75 / $3.75 | — | — | — | 75.1 |
| 42 | GPT-5.2 (high) (partial coverage) | OpenAI | 58.4 | 4 of 4 | $1.75 / $14.00 | 61.6 | 58.3 | 70.6 | 59.7 |
| 43 | GLM 5.2 (max) | Z.ai | 58.3 | 3 of 4 | $0.462 / $1.452 | 66.1 | 61.4 | 63.6 | — |
| 44 | Claude Opus 5 (medium) (partial coverage) | Anthropic | 58.2 | 1 of 4 | $5.00 / $25.00 | — | 74.3 | — | — |
| 45 | GPT 5.5 Pro Pre Release (xhigh) | OpenAI | 58.1 | 1 of 4 | — | — | — | 74.2 | — |
| 46 | GPT-5.5 Pro (xhigh) (partial coverage) | OpenAI | 57.9 | 1 of 4 | $30.00 / $180.00 | — | — | 73.5 | — |
| 47 | MiniMax M2.5 (high) (partial coverage) | MiniMax | 57.9 | 2 of 4 | $0.22 / $0.90 | 65.7 | 65.7 | — | — |
| 48 | GPT-5.6 Luna (xhigh) (partial coverage) | OpenAI | 57.9 | 4 of 4 | $0.10 / $0.60 | 60.8 | 62.8 | 63.5 | 59.9 |
| 49 | GPT-5.4 (partial coverage) | OpenAI | 57.8 | 3 of 4 | $2.50 / $15.00 | — | 57.0 | 59.4 | 72.6 |
| 50 | Grok 4.6 (xhigh) (partial coverage) | SpaceXAI | 57.8 | 2 of 4 | $2.00 / $6.00 | — | 62.7 | 68.1 | — |
| 51 | Claude Opus 4.5 (high 32K) (partial coverage) | Anthropic | 57.7 | 2 of 4 | $5.00 / $25.00 | — | 70.0 | 60.7 | — |
| 52 | Gemini 3.5 Flash (high) | Google | 57.7 | 4 of 4 | $1.50 / $9.00 | 47.6 | 56.1 | 72.8 | 69.6 |
| 53 | GPT-5.4 Pro (xhigh) | OpenAI | 57.6 | 1 of 4 | $30.00 / $180.00 | — | — | 72.5 | — |
| 54 | Claude Fable 5 (low) (partial coverage) | Anthropic | 57.5 | 2 of 4 | $10.00 / $50.00 | — | 66.1 | 63.8 | — |
| 55 | Qwen3.8 Max (xhigh) (partial coverage) | Qwen | 57.4 | 2 of 4 | $2.00 / $6.00 | — | 56.5 | 72.7 | — |
| 56 | Claude Sonnet 4.6 (partial coverage) | Anthropic | 57.3 | 4 of 4 | $3.00 / $15.00 | 52.1 | 70.2 | 59.8 | 61.3 |
| 57 | Dola Seed 2.0 Pro (partial coverage) | ByteDance | 57.3 | 3 of 4 | — | — | 66.1 | 57.9 | 62.0 |
| 58 | Grok 4.5 (high) (partial coverage) | SpaceXAI | 57.2 | 3 of 4 | $2.00 / $6.00 | 58.4 | 65.7 | 61.8 | — |
| 59 | Ernie 5.1 (partial coverage) | Baidu | 57.2 | 2 of 4 | — | — | 66.7 | 61.9 | — |
| 60 | Claude Opus 4.5 | Anthropic | 57.2 | 2 of 4 | $5.00 / $25.00 | — | 77.3 | 51.4 | — |
| 61 | Claude Opus 4.8 (high) (partial coverage) | Anthropic | 57.1 | 3 of 4 | $5.00 / $25.00 | 63.1 | 57.6 | 64.4 | — |
| 62 | GPT-5.6 Luna (max) (partial coverage) | OpenAI | 57.0 | 3 of 4 | $0.10 / $0.60 | 47.3 | 63.0 | 74.6 | — |
| 63 | DeepSeek V4 Flash 0423 (high) (partial coverage) | DeepSeek | 56.9 | 3 of 4 | $0.0643 / $0.1285 | 60.3 | 67.6 | 56.4 | — |
| 64 | Qwen3.5 Max Preview (partial coverage) | Qwen | 56.9 | 2 of 4 | — | — | 66.4 | 60.9 | — |
| 65 | Claude Opus 4.5 (high) (partial coverage) | Anthropic | 56.8 | 2 of 4 | $5.00 / $25.00 | 59.6 | 67.4 | — | — |
| 66 | Claude Fable 5 (medium) (partial coverage) | Anthropic | 56.7 | 1 of 4 | $10.00 / $50.00 | — | 69.9 | — | — |
| 67 | Grok 4.20 Multi Agent Beta 0309 (partial coverage) | xAI | 56.7 | 3 of 4 | — | — | 65.1 | 58.4 | 59.6 |
| 68 | Kimi K2.5 (thinking) (partial coverage) | MoonshotAI | 56.7 | 3 of 4 | $0.57 / $2.85 | — | 61.5 | 60.8 | 60.8 |
| 69 | Doubao-Seed-Code (partial coverage) | ByteDance | 56.6 | 1 of 4 | — | — | 69.7 | — | — |
| 70 | Claude Opus 4.6 (64K) | Anthropic | 56.6 | 1 of 4 | $5.00 / $25.00 | — | — | 69.6 | — |
| 71 | Gemini 3 Flash Preview (thinking minimal) (partial coverage) | Google | 56.6 | 3 of 4 | $0.50 / $3.00 | — | 55.4 | 58.6 | 68.7 |
| 72 | Qwen3.7 Max | Qwen | 56.5 | 3 of 4 | $1.475 / $4.425 | 45.6 | 66.8 | 69.9 | — |
| 73 | GLM-5.3 (partial coverage) | Z.ai | 56.5 | 1 of 4 | — | — | 69.1 | — | — |
| 74 | o4 Mini High (partial coverage) | OpenAI | 56.5 | 2 of 4 | $1.10 / $4.40 | — | 71.2 | 54.4 | — |
| 75 | Gemini 3 Pro Preview (partial coverage) | Google | 56.3 | 3 of 4 | — | 57.6 | 61.3 | 62.4 | — |
| 76 | GPT-5.6 Sol (high) | OpenAI | 56.3 | 1 of 4 | $5.00 / $30.00 | — | 68.6 | — | — |
| 77 | o3 (high) (partial coverage) | OpenAI | 56.3 | 2 of 4 | $2.00 / $8.00 | — | 69.9 | 54.9 | — |
| 78 | GPT 5.5 Pro Pre Release (high) (partial coverage) | OpenAI | 56.2 | 1 of 4 | — | — | — | 68.4 | — |
| 79 | Claude Opus 4.6 (32K) | Anthropic | 56.1 | 1 of 4 | $5.00 / $25.00 | — | — | 67.9 | — |
| 80 | Qwen3.5 397B A17B (partial coverage) | Qwen | 56.0 | 3 of 4 | $0.39 / $2.34 | — | 58.2 | 61.4 | 60.2 |
| 81 | Grok 4.20 Beta 0309 (reasoning) (partial coverage) | xAI | 56.0 | 3 of 4 | — | — | 58.3 | 60.1 | 61.4 |
| 82 | Gemini 3 Flash Preview | Google | 56.0 | 4 of 4 | $0.50 / $3.00 | 26.2 | 72.2 | 64.6 | 72.4 |
| 83 | Claude Opus 4.6 (high) (partial coverage) | Anthropic | 56.0 | 2 of 4 | $5.00 / $25.00 | — | 57.9 | 65.6 | — |
| 84 | Claude Opus 4.1 (thinking 16K) (partial coverage) | Anthropic | 55.9 | 2 of 4 | $15.00 / $75.00 | — | 66.0 | 57.4 | — |
| 85 | GPT 5.5 Instant (partial coverage) | OpenAI | 55.9 | 3 of 4 | — | — | 66.7 | 44.6 | 67.7 |
| 86 | Grok 4.6 (partial coverage) | SpaceXAI | 55.8 | 1 of 4 | $2.00 / $6.00 | — | 67.0 | — | — |
| 87 | Gemini 3.5 Flash (medium) (partial coverage) | Google | 55.8 | 4 of 4 | $1.50 / $9.00 | 33.2 | 66.0 | 62.2 | 72.8 |
| 88 | Kimi K2.6 | MoonshotAI | 55.7 | 4 of 4 | $0.5415 / $2.28 | 47.7 | 52.9 | 67.2 | 66.1 |
| 89 | GLM 5 (high) (partial coverage) | Z.ai | 55.7 | 2 of 4 | $0.60 / $1.92 | 61.6 | 60.8 | — | — |
| 90 | GPT-5.2 (medium) | OpenAI | 55.5 | 1 of 4 | $1.75 / $14.00 | — | — | 66.3 | — |
| 91 | Gemini 3.7 Flash (medium) (partial coverage) | Google | 55.5 | 1 of 4 | $0.375 / $1.875 | — | 66.3 | — | — |
| 92 | GPT-5.2 Pro (xhigh) (partial coverage) | OpenAI | 55.3 | 1 of 4 | $21.00 / $168.00 | — | — | 65.7 | — |
| 93 | GPT-5.4 (xhigh) (partial coverage) | OpenAI | 55.3 | 3 of 4 | $2.50 / $15.00 | 48.4 | 55.1 | 72.5 | — |
| 94 | Claude Sonnet 4.5 (high 32K) (partial coverage) | Anthropic | 55.2 | 2 of 4 | $3.00 / $15.00 | — | 61.5 | 58.9 | — |
| 95 | Qwen3.7 Plus (partial coverage) | Qwen | 55.1 | 4 of 4 | $0.32 / $1.28 | 38.4 | 64.7 | 65.9 | 61.6 |
| 96 | Claude Opus 4.7 (max) (partial coverage) | Anthropic | 55.1 | 3 of 4 | $5.00 / $25.00 | 46.5 | 65.5 | 63.3 | — |
| 97 | Gemini 3.1 Pro Preview Custom Tools (partial coverage) | Google | 55.1 | 1 of 4 | $2.00 / $12.00 | — | 65.1 | — | — |
| 98 | Kimi K3 (none) (partial coverage) | MoonshotAI | 55.1 | 1 of 4 | $3.00 / $15.00 | — | 65.0 | — | — |
| 99 | Claude Opus 5 (low) (partial coverage) | Anthropic | 55.0 | 2 of 4 | $5.00 / $25.00 | — | 60.2 | 59.5 | — |
| 100 | GPT-5 (high) | OpenAI | 55.0 | 3 of 4 | $1.25 / $10.00 | — | 56.9 | 65.5 | 52.4 |
| 101 | Claude Sonnet 4.5 (high) (partial coverage) | Anthropic | 55.0 | 2 of 4 | $3.00 / $15.00 | 60.1 | 59.5 | — | — |
| 102 | EXAONE 4.0 32B (partial coverage) | LG AI Research | 54.9 | 1 of 4 | — | — | 64.5 | — | — |
| 103 | Mimo v2 Pro (partial coverage) | Xiaomi | 54.9 | 2 of 4 | — | — | 61.2 | 58.2 | — |
| 104 | Grok 4.6 (medium) (partial coverage) | SpaceXAI | 54.9 | 1 of 4 | $2.00 / $6.00 | — | 64.4 | — | — |
| 105 | GPT-5 (medium) (partial coverage) | OpenAI | 54.9 | 3 of 4 | $1.25 / $10.00 | 52.7 | 57.7 | 63.6 | — |
| 106 | GPT-5.4 (medium) (partial coverage) | OpenAI | 54.8 | 2 of 4 | $2.50 / $15.00 | — | 57.4 | 61.6 | — |
| 107 | Seed 2.1 Pro Preview (partial coverage) | ByteDance | 54.8 | 1 of 4 | — | — | 64.0 | — | — |
| 108 | GPT-5.3-Codex (high) (partial coverage) | OpenAI | 54.8 | 1 of 4 | $1.75 / $14.00 | — | 64.0 | — | — |
| 109 | Qwen3 Max (partial coverage) | Qwen | 54.7 | 2 of 4 | $0.78 / $3.90 | — | 58.4 | 60.1 | — |
| 110 | Claude Opus 5 (partial coverage) | Anthropic | 54.7 | 1 of 4 | $5.00 / $25.00 | — | — | 63.8 | — |
| 111 | GLM 5.1 | Z.ai | 54.7 | 3 of 4 | $0.966 / $3.036 | 44.0 | 62.9 | 66.2 | — |
| 112 | GPT-5 (partial coverage) | OpenAI | 54.6 | 1 of 4 | $1.25 / $10.00 | — | 63.6 | — | — |
| 113 | Claude Opus 4.6 (max) (partial coverage) | Anthropic | 54.5 | 3 of 4 | $5.00 / $25.00 | 53.0 | 50.6 | 68.7 | — |
| 114 | DeepSeek V4 Pro (max) (partial coverage) | DeepSeek | 54.5 | 2 of 4 | $1.168 / $2.336 | — | 67.5 | 50.2 | — |
| 115 | Kimi K2.5 (high) (partial coverage) | MoonshotAI | 54.5 | 2 of 4 | $0.57 / $2.85 | 59.4 | 58.3 | — | — |
| 116 | OpenReasoning Nemotron 32B (partial coverage) | NVIDIA | 54.5 | 1 of 4 | — | — | 63.2 | — | — |
| 117 | GPT-5.2 (xhigh) (partial coverage) | OpenAI | 54.5 | 3 of 4 | $1.75 / $14.00 | 43.8 | 59.4 | 68.8 | — |
| 118 | Claude Sonnet 5 (max) (partial coverage) | Anthropic | 54.4 | 2 of 4 | $2.00 / $10.00 | — | 58.7 | 58.6 | — |
| 119 | Ernie 5.0 0110 (partial coverage) | Baidu | 54.3 | 2 of 4 | — | — | 61.2 | 55.8 | — |
| 120 | GLM 5.2 (partial coverage) | Z.ai | 54.3 | 2 of 4 | $0.462 / $1.452 | 54.1 | 62.9 | — | — |
| 121 | Claude Sonnet 5 (xhigh) (partial coverage) | Anthropic | 54.2 | 2 of 4 | $2.00 / $10.00 | — | 56.0 | 60.7 | — |
| 122 | Grok 4.1 (partial coverage) | xAI | 54.2 | 2 of 4 | — | — | 62.1 | 54.4 | — |
| 123 | Claude Opus 4.8 (xhigh) (partial coverage) | Anthropic | 54.1 | 1 of 4 | $5.00 / $25.00 | — | 62.2 | — | — |
| 124 | GPT-5.6 Luna (high) | OpenAI | 54.1 | 1 of 4 | $0.10 / $0.60 | — | 62.1 | — | — |
| 125 | GPT-5.3 Chat (partial coverage) | OpenAI | 54.1 | 2 of 4 | — | — | 62.5 | 53.6 | — |
| 126 | SWE-1.7 (none) (partial coverage) | Cognition | 53.9 | 1 of 4 | — | — | 61.5 | — | — |
| 127 | DeepSeek V3.2 (high) (partial coverage) | DeepSeek | 53.9 | 2 of 4 | $0.269 / $0.40 | 57.9 | 57.6 | — | — |
| 128 | GLM 5 | Z.ai | 53.8 | 2 of 4 | $0.60 / $1.92 | — | 64.1 | 50.9 | — |
| 129 | GPT-5.1 (high) (partial coverage) | OpenAI | 53.8 | 4 of 4 | $1.25 / $10.00 | 35.7 | 59.9 | 62.3 | 64.7 |
| 130 | Gemini 2.5 Pro Preview 05-06 (partial coverage) | Google | 53.7 | 1 of 4 | $1.25 / $10.00 | — | — | 60.9 | — |
| 131 | Gemma 4 26B A4B  (partial coverage) | Google | 53.7 | 3 of 4 | $0.12 / $0.40 | — | 53.3 | 60.2 | 54.6 |
| 132 | Grok 4.5 (medium) (partial coverage) | SpaceXAI | 53.6 | 1 of 4 | $2.00 / $6.00 | — | 60.7 | — | — |
| 133 | Gemini 3 Pro Preview (high) (partial coverage) | Google | 53.6 | 2 of 4 | — | 57.1 | 57.0 | — | — |
| 134 | OpenCodeReasoning Nemotron 1.1 32B (partial coverage) | NVIDIA | 53.6 | 1 of 4 | — | — | 60.5 | — | — |
| 135 | DeepSeek V4 Flash 0731 (max) | DeepSeek | 53.6 | 1 of 4 | $0.14 / $0.28 | — | — | 60.4 | — |
| 136 | DeepSeek V3.2 Exp (thinking) (partial coverage) | DeepSeek | 53.5 | 2 of 4 | $0.27 / $0.41 | — | 59.0 | 54.6 | — |
| 137 | Qwen3.6 Plus | Qwen | 53.4 | 2 of 4 | $0.325 / $1.95 | — | 48.1 | 65.2 | — |
| 138 | o4 Mini (medium) (partial coverage) | OpenAI | 53.4 | 2 of 4 | $1.10 / $4.40 | — | 68.5 | 44.7 | — |
| 139 | GPT-5 Pro (high) | OpenAI | 53.3 | 1 of 4 | $15.00 / $120.00 | — | — | 59.6 | — |
| 140 | Kimi K3 (high) (partial coverage) | MoonshotAI | 53.3 | 1 of 4 | $3.00 / $15.00 | — | — | 59.5 | — |
| 141 | GPT-5.6 Sol (medium) (partial coverage) | OpenAI | 53.3 | 1 of 4 | $5.00 / $30.00 | — | 59.5 | — | — |
| 142 | MiMo-V2.5 (partial coverage) | Xiaomi | 53.3 | 3 of 4 | $0.14 / $0.28 | — | 60.6 | 56.6 | 48.8 |
| 143 | Gemini 3.6 Flash (high) | Google | 53.2 | 3 of 4 | $0.75 / $3.75 | 40.2 | 59.1 | 66.6 | — |
| 144 | Kimi K2.5 Instant (partial coverage) | Moonshot AI | 53.1 | 3 of 4 | — | — | 60.4 | 56.1 | 49.0 |
| 145 | MiMo-V2.5-Pro (partial coverage) | Xiaomi | 53.1 | 3 of 4 | $0.435 / $0.87 | 35.5 | 67.6 | 62.1 | — |
| 146 | GPT-5.4 Mini (high) | OpenAI | 53.0 | 3 of 4 | $0.75 / $4.50 | — | 50.1 | 57.6 | 57.1 |
| 147 | GPT-5.6 Sol (partial coverage) | OpenAI | 53.0 | 1 of 4 | $5.00 / $30.00 | 58.7 | — | — | — |
| 148 | Muse Glimmer 30B (partial coverage) | Meta | 52.9 | 2 of 4 | $0.35 / $1.50 | — | 52.9 | 58.4 | — |
| 149 | Hy3 (partial coverage) | Tencent | 52.8 | 3 of 4 | $0.132 / $0.528 | 32.5 | 67.8 | 63.2 | — |
| 150 | Claude Haiku 4.5 (high) (partial coverage) | Anthropic | 52.7 | 2 of 4 | $1.00 / $5.00 | 54.9 | 55.7 | — | — |
| 151 | Qwen3.6 27B (partial coverage) | Qwen | 52.7 | 1 of 4 | $0.60 / $3.60 | — | — | 57.9 | — |
| 152 | Kimi K2 0905 (partial coverage) | MoonshotAI | 52.6 | 2 of 4 | $0.60 / $2.50 | — | 58.7 | 51.5 | — |
| 153 | Inkling Small (partial coverage) | Thinking Machines | 52.6 | 1 of 4 | $0.45 / $1.20 | 57.6 | — | — | — |
| 154 | Qwen3 235B A22B Instruct 2507 (partial coverage) | Qwen | 52.6 | 2 of 4 | $0.09 / $0.55 | — | 58.3 | 51.8 | — |
| 155 | GPT-5.2 (partial coverage) | OpenAI | 52.6 | 4 of 4 | $1.75 / $14.00 | 56.4 | 57.4 | 54.9 | 46.5 |
| 156 | R1 0528 (partial coverage) | DeepSeek | 52.6 | 2 of 4 | $0.50 / $2.15 | — | 63.7 | 46.3 | — |
| 157 | Mimo v2 Omni (partial coverage) | Xiaomi | 52.5 | 3 of 4 | — | — | 60.8 | 55.2 | 46.5 |
| 158 | LongCat Flash Chat (partial coverage) | Meituan | 52.5 | 2 of 4 | — | — | 58.7 | 50.9 | — |
| 159 | o3 (partial coverage) | OpenAI | 52.4 | 4 of 4 | $2.00 / $8.00 | 48.3 | 61.5 | 57.6 | 46.8 |
| 160 | Gemini 3.5 Flash (low) (partial coverage) | Google | 52.4 | 1 of 4 | $1.50 / $9.00 | — | — | 56.8 | — |
| 161 | gpt-oss-120b (high) (partial coverage) | OpenAI | 52.4 | 1 of 4 | $0.03 / $0.17 | — | — | 56.8 | — |
| 162 | DeepSeek V3.1 Terminus (thinking) (partial coverage) | DeepSeek | 52.3 | 1 of 4 | $0.27 / $0.95 | — | 56.6 | — | — |
| 163 | GLM 5V Turbo (partial coverage) | Z.ai | 52.3 | 3 of 4 | $1.20 / $4.00 | — | 57.6 | 57.2 | 46.4 |
| 164 | GPT-5.1-Codex (medium) (partial coverage) | OpenAI | 52.3 | 2 of 4 | $1.25 / $10.00 | 53.8 | 55.1 | — | — |
| 165 | Chatgpt 4o (partial coverage) | OpenAI | 52.2 | 3 of 4 | — | — | 57.8 | 48.6 | 54.5 |
| 166 | GPT-5.6 Sol (low) (partial coverage) | OpenAI | 52.2 | 2 of 4 | $5.00 / $30.00 | — | 47.0 | 61.6 | — |
| 167 | Claude Opus 4.7 (xhigh) (partial coverage) | Anthropic | 52.1 | 2 of 4 | $5.00 / $25.00 | — | 51.9 | 56.3 | — |
| 168 | Claude Opus 4.8 (low) (partial coverage) | Anthropic | 52.1 | 2 of 4 | $5.00 / $25.00 | — | 44.3 | 63.8 | — |
| 169 | Gemini 3.1 Flash Lite Preview (partial coverage) | Google | 52.1 | 3 of 4 | $0.25 / $1.50 | — | 44.5 | 55.6 | 60.0 |
| 170 | Ernie 5.0 Preview 1203 (partial coverage) | Baidu | 52.1 | 2 of 4 | — | — | 58.2 | 49.9 | — |
| 171 | GPT-5.2-Codex (partial coverage) | OpenAI | 52.0 | 2 of 4 | $1.75 / $14.00 | 61.6 | 46.3 | — | — |
| 172 | Grok 4.5 (low) (partial coverage) | SpaceXAI | 52.0 | 1 of 4 | $2.00 / $6.00 | — | 55.8 | — | — |
| 173 | GPT-5 Mini (medium) (partial coverage) | OpenAI | 52.0 | 3 of 4 | $0.25 / $2.00 | 49.0 | 54.1 | 56.8 | — |
| 174 | o1 (medium) (partial coverage) | OpenAI | 52.0 | 1 of 4 | $15.00 / $60.00 | — | — | 55.7 | — |
| 175 | GPT-5.2 (low) | OpenAI | 51.9 | 1 of 4 | $1.75 / $14.00 | — | — | 55.5 | — |
| 176 | GPT-5.1 (medium) (partial coverage) | OpenAI | 51.9 | 3 of 4 | $1.25 / $10.00 | 53.8 | 51.9 | 53.7 | — |
| 177 | Qwen3.6 35B A3B (partial coverage) | Qwen | 51.8 | 1 of 4 | $0.15 / $1.00 | — | — | 55.3 | — |
| 178 | Qwen3.7 Flash (partial coverage) | Qwen | 51.8 | 1 of 4 | $0.03 / $0.13 | — | — | 55.3 | — |
| 179 | XBai o4 (medium) (partial coverage) | MetaStoneTec | 51.8 | 1 of 4 | — | — | 55.2 | — | — |
| 180 | Kimi K2 Thinking Turbo (partial coverage) | Moonshot AI | 51.8 | 2 of 4 | — | — | 50.4 | 56.5 | — |
| 181 | Qwen3.5 Plus | Qwen | 51.7 | 1 of 4 | $0.30 / $1.80 | — | — | 54.8 | — |
| 182 | MiniMax M3 (partial coverage) | MiniMax | 51.7 | 4 of 4 | $0.30 / $1.20 | 37.5 | 64.8 | 53.6 | 53.7 |
| 183 | Claude Opus 4.5 (128K) (partial coverage) | Anthropic | 51.6 | 1 of 4 | $5.00 / $25.00 | — | 54.5 | — | — |
| 184 | Claude Sonnet 4.6 (32K) (partial coverage) | Anthropic | 51.5 | 1 of 4 | $3.00 / $15.00 | — | — | 54.4 | — |
| 185 | GPT-5.3-Codex (partial coverage) | OpenAI | 51.5 | 1 of 4 | $1.75 / $14.00 | — | 54.4 | — | — |
| 186 | GPT 5 Chat (partial coverage) | OpenAI | 51.5 | 3 of 4 | — | — | 56.5 | 50.3 | 50.6 |
| 187 | DeepSeek V4 Pro (xhigh) (partial coverage) | DeepSeek | 51.5 | 1 of 4 | $1.168 / $2.336 | — | 54.4 | — | — |
| 188 | Grok 4.1 (thinking) (partial coverage) | xAI | 51.5 | 2 of 4 | — | — | 48.9 | 56.9 | — |
| 189 | DeepSeek V3.1 (thinking) (partial coverage) | DeepSeek | 51.5 | 2 of 4 | $0.25 / $0.95 | — | 55.0 | 50.8 | — |
| 190 | o3 (medium) (partial coverage) | OpenAI | 51.4 | 2 of 4 | $2.00 / $8.00 | — | 52.6 | 52.6 | — |
| 191 | Claude Opus 4.8 (none) (partial coverage) | Anthropic | 51.3 | 1 of 4 | $5.00 / $25.00 | — | — | 53.5 | — |
| 192 | GPT-5.4 (low) (partial coverage) | OpenAI | 51.3 | 1 of 4 | $2.50 / $15.00 | — | — | 53.5 | — |
| 193 | Qwen3 Next 80B A3B Instruct (partial coverage) | Qwen | 51.2 | 2 of 4 | $0.10 / $1.10 | — | 53.4 | 51.3 | — |
| 194 | Kimi K2.7 Code (partial coverage) | MoonshotAI | 51.2 | 3 of 4 | $0.71 / $3.50 | 52.7 | 47.8 | 55.3 | — |
| 195 | GPT-5.5 (medium) (partial coverage) | OpenAI | 51.1 | 1 of 4 | $5.00 / $30.00 | — | 53.2 | — | — |
| 196 | DeepSeek V3.1 (partial coverage) | DeepSeek | 51.1 | 2 of 4 | $0.25 / $0.95 | — | 53.6 | 50.6 | — |
| 197 | R1 (partial coverage) | DeepSeek | 51.1 | 2 of 4 | $0.70 / $2.50 | — | 53.1 | 51.1 | — |
| 198 | Inkling Small (xhigh) | Thinking Machines | 51.1 | 1 of 4 | $0.45 / $1.20 | — | — | 53.0 | — |
| 199 | Step 3.5 Flash (partial coverage) | StepFun | 50.9 | 2 of 4 | $0.10 / $0.30 | — | 54.2 | 49.4 | — |
| 200 | Gemini 3.7 Flash (low) (partial coverage) | Google | 50.9 | 1 of 4 | $0.375 / $1.875 | — | 52.6 | — | — |
| 201 | Grok 4.20 0309 (reasoning) | xAI | 50.9 | 1 of 4 | — | — | — | 52.6 | — |
| 202 | Qwen3 VL 235B A22B Instruct (partial coverage) | Qwen | 50.9 | 3 of 4 | $0.26 / $1.04 | — | 57.1 | 49.6 | 47.7 |
| 203 | o1 (high) | OpenAI | 50.9 | 1 of 4 | $15.00 / $60.00 | — | — | 52.4 | — |
| 204 | Gemini 2.5 Pro (partial coverage) | Google | 50.9 | 4 of 4 | $1.25 / $10.00 | 43.1 | 50.6 | 47.7 | 63.8 |
| 205 | Qwen3.5 397B A17B (none) (partial coverage) | Qwen | 50.9 | 1 of 4 | $0.39 / $2.34 | — | — | 52.3 | — |
| 206 | GPT-5.6 Terra (high) (partial coverage) | OpenAI | 50.8 | 1 of 4 | $1.00 / $6.00 | — | 52.2 | — | — |
| 207 | Claude Opus 4 (thinking 16K) (partial coverage) | Anthropic | 50.8 | 3 of 4 | $15.00 / $75.00 | — | 63.5 | 52.2 | 38.1 |
| 208 | o3 Mini (medium) | OpenAI | 50.8 | 1 of 4 | $1.10 / $4.40 | — | — | 52.1 | — |
| 209 | GPT-5.5 (low) (partial coverage) | OpenAI | 50.8 | 2 of 4 | $5.00 / $30.00 | — | 49.3 | 53.5 | — |
| 210 | Claude Sonnet 4.5 (partial coverage) | Anthropic | 50.7 | 3 of 4 | $3.00 / $15.00 | 58.6 | 54.5 | 40.3 | — |
| 211 | Hunyuan Hy3 Preview (partial coverage) | Tencent | 50.7 | 2 of 4 | — | — | 49.4 | 53.2 | — |
| 212 | DeepSeek V4 Pro (partial coverage) | DeepSeek | 50.7 | 3 of 4 | $1.168 / $2.336 | 41.7 | 53.9 | 57.5 | — |
| 213 | DeepSeek V3.2 (thinking) (partial coverage) | DeepSeek | 50.6 | 3 of 4 | $0.269 / $0.40 | 49.7 | 50.1 | 53.1 | — |
| 214 | Gemini 3.6 Flash (medium) (partial coverage) | Google | 50.6 | 1 of 4 | $0.75 / $3.75 | — | 51.4 | — | — |
| 215 | Claude Opus 4.5 (16K) | Anthropic | 50.6 | 1 of 4 | $5.00 / $25.00 | — | — | 51.4 | — |
| 216 | DeepSeek V4 Flash 0423 (max) (partial coverage) | DeepSeek | 50.5 | 1 of 4 | $0.0643 / $0.1285 | — | 51.1 | — | — |
| 217 | Gemini 3.5 Flash (minimal) (partial coverage) | Google | 50.5 | 1 of 4 | $1.50 / $9.00 | — | — | 51.1 | — |
| 218 | Gemini 3.6 Flash (minimal) (partial coverage) | Google | 50.5 | 1 of 4 | $0.75 / $3.75 | — | — | 51.1 | — |
| 219 | GPT 4.5 Preview (partial coverage) | OpenAI | 50.4 | 3 of 4 | — | — | 55.5 | 43.9 | 52.5 |
| 220 | Muse Spark 1.1 (xhigh) (partial coverage) | Meta | 50.4 | 2 of 4 | $1.25 / $4.25 | 50.1 | 51.1 | — | — |
| 221 | o4 Mini (low) (partial coverage) | OpenAI | 50.3 | 2 of 4 | $1.10 / $4.40 | — | 57.2 | 43.9 | — |
| 222 | Claude Sonnet 4.6 (16K) (partial coverage) | Anthropic | 50.3 | 1 of 4 | $3.00 / $15.00 | — | — | 50.6 | — |
| 223 | GPT-5.1 (partial coverage) | OpenAI | 50.3 | 3 of 4 | $1.25 / $10.00 | — | 50.7 | 52.5 | 47.9 |
| 224 | Gemini 3.5 Flash Lite (partial coverage) | Google | 50.2 | 4 of 4 | $0.30 / $2.50 | 21.5 | 63.4 | 52.9 | 63.2 |
| 225 | Kimi K2p5 (partial coverage) | Moonshot AI | 50.1 | 2 of 4 | — | 42.6 | — | 57.6 | — |
| 226 | Kimi K2.5 | MoonshotAI | 50.1 | 1 of 4 | $0.57 / $2.85 | — | 50.1 | — | — |
| 227 | Kimi K2 0711 (partial coverage) | MoonshotAI | 50.1 | 2 of 4 | $0.57 / $2.30 | — | 54.6 | 45.6 | — |
| 228 | Qwen3.7 Flash (none) (partial coverage) | Qwen | 50.1 | 1 of 4 | $0.03 / $0.13 | — | — | 50.0 | — |
| 229 | Qwen3.5-122B-A10B (partial coverage) | Qwen | 50.1 | 3 of 4 | $0.29 / $2.40 | — | 49.4 | 52.6 | 48.1 |
| 230 | Claude Opus 4 (thinking) (partial coverage) | Anthropic | 50.0 | 1 of 4 | $15.00 / $75.00 | — | 49.9 | — | — |
| 231 | Qwen3 235B A22B (nothinking) (partial coverage) | Qwen | 50.0 | 2 of 4 | $0.455 / $1.82 | — | 53.2 | 46.5 | — |
| 232 | Kimi K2.7 Code (none) (partial coverage) | MoonshotAI | 50.0 | 1 of 4 | $0.71 / $3.50 | — | 49.7 | — | — |
| 233 | Claude Sonnet 4 (thinking 32K) (partial coverage) | Anthropic | 50.0 | 3 of 4 | $3.00 / $15.00 | — | 58.6 | 48.1 | 43.0 |
| 234 | DeepSeek V4 Flash 0423 (partial coverage) | DeepSeek | 49.9 | 3 of 4 | $0.0643 / $0.1285 | 35.5 | 60.4 | 53.3 | — |
| 235 | o3 Mini High (partial coverage) | OpenAI | 49.9 | 2 of 4 | $1.10 / $4.40 | — | 56.4 | 42.8 | — |
| 236 | Nemotron 3.5 Lightning 30B A3B (partial coverage) | NVIDIA | 49.8 | 1 of 4 | — | — | 49.2 | — | — |
| 237 | Claude Sonnet 4.5 (59K) (partial coverage) | Anthropic | 49.8 | 1 of 4 | $3.00 / $15.00 | — | — | 49.1 | — |
| 238 | DeepSeek V3.1 Terminus (partial coverage) | DeepSeek | 49.8 | 2 of 4 | $0.27 / $0.95 | — | 52.1 | 46.6 | — |
| 239 | Gemma 4 31B (minimal) (partial coverage) | Google | 49.7 | 1 of 4 | $0.10 / $0.34 | — | — | 48.9 | — |
| 240 | Qwen3 235B A22B Thinking 2507 (partial coverage) | Qwen | 49.6 | 2 of 4 | $0.23 / $2.30 | — | 52.5 | 45.8 | — |
| 241 | Claude Opus 4.1 | Anthropic | 49.6 | 2 of 4 | $15.00 / $75.00 | — | 53.9 | 44.3 | — |
| 242 | Inkling (xhigh) (partial coverage) | Thinking Machines | 49.6 | 2 of 4 | $0.95 / $4.05 | 51.8 | — | 46.4 | — |
| 243 | Claude Sonnet 5 (partial coverage) | Anthropic | 49.6 | 1 of 4 | $2.00 / $10.00 | — | 48.6 | — | — |
| 244 | Claude Sonnet 4 (thinking) (partial coverage) | Anthropic | 49.6 | 1 of 4 | $3.00 / $15.00 | — | 48.5 | — | — |
| 245 | DeepSeek V3.2 Exp (partial coverage) | DeepSeek | 49.5 | 2 of 4 | $0.27 / $0.41 | — | 46.7 | 51.2 | — |
| 246 | Claude Opus 4.7 (low) (partial coverage) | Anthropic | 49.5 | 1 of 4 | $5.00 / $25.00 | — | 48.4 | — | — |
| 247 | MiniMax M2.5 (partial coverage) | MiniMax | 49.5 | 2 of 4 | $0.22 / $0.90 | — | 51.0 | 46.9 | — |
| 248 | Composer 2.5 (partial coverage) | Cursor | 49.5 | 1 of 4 | — | — | 48.3 | — | — |
| 249 | Claude Sonnet 4 (32K) (partial coverage) | Anthropic | 49.5 | 1 of 4 | $3.00 / $15.00 | — | — | 48.1 | — |
| 250 | Claude Sonnet 4.5 (16K) (partial coverage) | Anthropic | 49.5 | 1 of 4 | $3.00 / $15.00 | — | — | 48.1 | — |
| 251 | Kimi K2 Thinking (partial coverage) | MoonshotAI | 49.4 | 3 of 4 | $0.60 / $2.50 | 51.2 | 53.4 | 42.3 | — |
| 252 | Ernie 5.0 Preview 1022 (partial coverage) | Baidu | 49.4 | 2 of 4 | — | — | 50.3 | 47.1 | — |
| 253 | Claude Opus 4.5 (32K) | Anthropic | 49.4 | 1 of 4 | $5.00 / $25.00 | — | — | 47.9 | — |
| 254 | Gemini 2.5 Pro Preview 06-05 (partial coverage) | Google | 49.4 | 1 of 4 | $1.25 / $10.00 | — | — | 47.9 | — |
| 255 | Devstral Small 2512 (partial coverage) | Mistral AI | 49.3 | 2 of 4 | — | 47.5 | 49.6 | — | — |
| 256 | Qwen3 Coder 30B A3B Instruct (partial coverage) | Qwen | 49.3 | 1 of 4 | $0.07 / $0.28 | — | 47.7 | — | — |
| 257 | Claude Sonnet 4.5 (32K) | Anthropic | 49.3 | 1 of 4 | $3.00 / $15.00 | — | — | 47.5 | — |
| 258 | Gemma 4 31B (partial coverage) | Google | 49.2 | 4 of 4 | $0.10 / $0.34 | 22.7 | 56.0 | 61.2 | 55.3 |
| 259 | Claude Haiku 4.5 (32K) | Anthropic | 49.2 | 1 of 4 | $1.00 / $5.00 | — | — | 47.4 | — |
| 260 | Qwen3 30B A3B Instruct 2507 (partial coverage) | Qwen | 49.2 | 2 of 4 | $0.0482 / $0.1931 | — | 52.3 | 44.3 | — |
| 261 | GPT-5 Mini (high) (partial coverage) | OpenAI | 49.2 | 3 of 4 | $0.25 / $2.00 | — | 50.1 | 57.9 | 37.6 |
| 262 | Kimi K3 (low) (partial coverage) | MoonshotAI | 49.1 | 1 of 4 | $3.00 / $15.00 | — | — | 47.0 | — |
| 263 | GPT-5.4 Nano (low) (partial coverage) | OpenAI | 49.1 | 1 of 4 | $0.20 / $1.25 | — | — | 47.0 | — |
| 264 | GPT-5.6 Sol (none) (partial coverage) | OpenAI | 49.1 | 1 of 4 | $5.00 / $30.00 | — | — | 47.0 | — |
| 265 | Qwen3.6 35B A3B (none) (partial coverage) | Qwen | 49.1 | 1 of 4 | $0.15 / $1.00 | — | — | 47.0 | — |
| 266 | Devstral Small 2505 (partial coverage) | Mistral AI | 49.1 | 1 of 4 | — | — | 46.9 | — | — |
| 267 | Laguna M.1 (partial coverage) | Poolside | 49.0 | 1 of 4 | — | — | 46.9 | — | — |
| 268 | Gemini 3.6 Flash (low) (partial coverage) | Google | 49.0 | 2 of 4 | $0.75 / $3.75 | — | 43.6 | 52.3 | — |
| 269 | GPT-5.1 (low) (partial coverage) | OpenAI | 49.0 | 1 of 4 | $1.25 / $10.00 | — | — | 46.7 | — |
| 270 | Composer 2.5 (none) (partial coverage) | Cursor | 49.0 | 1 of 4 | — | — | 46.6 | — | — |
| 271 | GLM 4.5 Air (partial coverage) | Z.ai | 49.0 | 2 of 4 | $0.13 / $0.85 | — | 49.6 | 45.9 | — |
| 272 | Claude Sonnet 4.6 (medium) (partial coverage) | Anthropic | 48.9 | 2 of 4 | $3.00 / $15.00 | — | 43.1 | 52.3 | — |
| 273 | Qwen3 32B (partial coverage) | Qwen | 48.9 | 2 of 4 | $0.08 / $0.28 | — | 47.7 | 47.6 | — |
| 274 | Qwen3 Next 80B A3B Thinking (partial coverage) | Qwen | 48.9 | 2 of 4 | $0.15 / $1.20 | — | 49.1 | 46.2 | — |
| 275 | MiniMax M2.1 (partial coverage) | MiniMax | 48.9 | 2 of 4 | $0.30 / $1.20 | — | 49.2 | 46.1 | — |
| 276 | Hunyuan T1 (partial coverage) | Tencent | 48.8 | 2 of 4 | — | — | 47.2 | 47.9 | — |
| 277 | Qwen3.6 27B (none) (partial coverage) | Qwen | 48.8 | 1 of 4 | $0.60 / $3.60 | — | — | 46.2 | — |
| 278 | GPT-5.4 Nano (high) (partial coverage) | OpenAI | 48.8 | 3 of 4 | $0.20 / $1.25 | — | 56.0 | 52.9 | 34.7 |
| 279 | Inkling (partial coverage) | Thinking Machines | 48.7 | 3 of 4 | $0.95 / $4.05 | 32.2 | 48.0 | 62.9 | — |
| 280 | Trinity Large Preview (partial coverage) | Arcee AI | 48.7 | 2 of 4 | — | — | 52.7 | 41.8 | — |
| 281 | Grok 3 Mini Beta (low) | xAI | 48.7 | 1 of 4 | — | — | — | 45.7 | — |
| 282 | GPT-5.1-Codex (partial coverage) | OpenAI | 48.6 | 1 of 4 | $1.25 / $10.00 | — | 45.7 | — | — |
| 283 | Qwen3.5-27B (partial coverage) | Qwen | 48.6 | 3 of 4 | $0.195 / $1.56 | — | 47.9 | 54.2 | 40.8 |
| 284 | Nova Premier 1.0 (partial coverage) | Amazon | 48.6 | 1 of 4 | $2.50 / $12.50 | — | 45.6 | — | — |
| 285 | Claude Sonnet 4.6 (max) (partial coverage) | Anthropic | 48.5 | 2 of 4 | $3.00 / $15.00 | — | 45.8 | 48.1 | — |
| 286 | Qwen3.6 Flash | Qwen | 48.5 | 1 of 4 | $0.1875 / $1.125 | — | — | 45.2 | — |
| 287 | Grok 4.6 (low) (partial coverage) | SpaceXAI | 48.5 | 1 of 4 | $2.00 / $6.00 | — | 45.2 | — | — |
| 288 | GPT-5.2 (none) (partial coverage) | OpenAI | 48.4 | 1 of 4 | $1.75 / $14.00 | — | — | 45.0 | — |
| 289 | Claude Opus 4.6 (medium) (partial coverage) | Anthropic | 48.4 | 1 of 4 | $5.00 / $25.00 | — | 44.9 | — | — |
| 290 | Grok 3 Mini Beta (high) (partial coverage) | xAI | 48.4 | 2 of 4 | — | — | 50.6 | 42.6 | — |
| 291 | Claude 3.7 Sonnet (32K) | Anthropic | 48.3 | 1 of 4 | — | — | — | 44.7 | — |
| 292 | INTELLECT-3 (partial coverage) | Prime Intellect | 48.3 | 2 of 4 | — | — | 48.1 | 44.9 | — |
| 293 | Claude Opus 4 (16K) (partial coverage) | Anthropic | 48.3 | 1 of 4 | $15.00 / $75.00 | — | — | 44.6 | — |
| 294 | Gemini 3.5 Flash Lite (low) (partial coverage) | Google | 48.3 | 1 of 4 | $0.30 / $2.50 | — | — | 44.6 | — |
| 295 | Claude Sonnet 5 (medium) (partial coverage) | Anthropic | 48.3 | 1 of 4 | $2.00 / $10.00 | — | 44.6 | — | — |
| 296 | o1 (partial coverage) | OpenAI | 48.3 | 3 of 4 | $15.00 / $60.00 | — | 50.0 | 43.9 | 47.2 |
| 297 | Llama 3.3 Nemotron Super 49B v1.5 (partial coverage) | NVIDIA | 48.3 | 2 of 4 | — | — | 47.5 | 45.3 | — |
| 298 | GLM 4.7 Flash (partial coverage) | Z.ai | 48.2 | 2 of 4 | $0.06 / $0.40 | — | 49.5 | 42.9 | — |
| 299 | Laguna XS.2 (partial coverage) | Poolside | 48.1 | 1 of 4 | — | — | 44.2 | — | — |
| 300 | GPT-5.4 (none) (partial coverage) | OpenAI | 48.1 | 1 of 4 | $2.50 / $15.00 | — | — | 44.0 | — |
| 301 | GPT-5.5 (none) (partial coverage) | OpenAI | 48.1 | 1 of 4 | $5.00 / $30.00 | — | — | 44.0 | — |
| 302 | GPT-5.6 Terra (low) (partial coverage) | OpenAI | 48.1 | 2 of 4 | $1.00 / $6.00 | — | 35.3 | 56.8 | — |
| 303 | MiniMax M1 (partial coverage) | MiniMax | 48.1 | 2 of 4 | $0.55 / $2.20 | — | 48.7 | 43.3 | — |
| 304 | Devstral Small 2507 (partial coverage) | Mistral AI | 48.1 | 1 of 4 | — | — | 43.9 | — | — |
| 305 | Nemotron 3 Super (partial coverage) | NVIDIA | 48.0 | 2 of 4 | $0.085 / $0.40 | — | 47.9 | 44.0 | — |
| 306 | GLM 4.7 (partial coverage) | Z.ai | 48.0 | 3 of 4 | $0.40 / $1.75 | 39.2 | 58.7 | 42.1 | — |
| 307 | DeepSeek V3.2 | DeepSeek | 48.0 | 2 of 4 | $0.269 / $0.40 | — | 40.6 | 51.2 | — |
| 308 | Qwen3 VL 235B A22B Thinking (partial coverage) | Qwen | 48.0 | 3 of 4 | $0.40 / $4.00 | — | 54.6 | 48.5 | 36.7 |
| 309 | DeepSeek V4 Flash 0423 (xhigh) (partial coverage) | DeepSeek | 48.0 | 1 of 4 | $0.0643 / $0.1285 | — | 43.8 | — | — |
| 310 | o3 Mini (low) (partial coverage) | OpenAI | 48.0 | 2 of 4 | $1.10 / $4.40 | — | 51.2 | 40.6 | — |
| 311 | Mercury (partial coverage) | Inception | 48.0 | 1 of 4 | — | — | 43.8 | — | — |
| 312 | Claude Opus 4.1 (27K) | Anthropic | 48.0 | 1 of 4 | $15.00 / $75.00 | — | — | 43.7 | — |
| 313 | Grok 4.3 (partial coverage) | SpaceXAI | 48.0 | 4 of 4 | $1.25 / $2.50 | 27.4 | 52.7 | 51.9 | 55.6 |
| 314 | GPT-5 Mini (minimal) (partial coverage) | OpenAI | 47.9 | 1 of 4 | $0.25 / $2.00 | — | — | 43.6 | — |
| 315 | Ernie 5.0 Preview 1220 | Baidu | 47.9 | 1 of 4 | — | — | — | — | 43.4 |
| 316 | o3 (low) (partial coverage) | OpenAI | 47.9 | 1 of 4 | $2.00 / $8.00 | — | — | 43.4 | — |
| 317 | Llama 3.3 Nemotron Super 49B v1 (partial coverage) | NVIDIA | 47.9 | 1 of 4 | — | — | 43.3 | — | — |
| 318 | gpt-oss-20b (high) (partial coverage) | OpenAI | 47.9 | 1 of 4 | $0.03 / $0.13 | — | — | 43.3 | — |
| 319 | Llama 3.1 Nemotron Ultra 253B v1 (partial coverage) | NVIDIA | 47.8 | 2 of 4 | — | — | 46.5 | 44.5 | — |
| 320 | Claude 3.7 Sonnet (16K) | Anthropic | 47.8 | 1 of 4 | — | — | — | 43.0 | — |
| 321 | Grok 3 Beta (partial coverage) | xAI | 47.8 | 2 of 4 | — | — | 52.8 | 37.9 | — |
| 322 | o3 Mini (partial coverage) | OpenAI | 47.7 | 2 of 4 | $1.10 / $4.40 | — | 45.8 | 44.8 | — |
| 323 | Claude Sonnet 4 (16K) (partial coverage) | Anthropic | 47.7 | 1 of 4 | $3.00 / $15.00 | — | — | 42.9 | — |
| 324 | GPT-5.6 Terra (none) (partial coverage) | OpenAI | 47.7 | 1 of 4 | $1.00 / $6.00 | — | — | 42.9 | — |
| 325 | o1 (low) (partial coverage) | OpenAI | 47.7 | 1 of 4 | $15.00 / $60.00 | — | — | 42.9 | — |
| 326 | Qwen3.7 Plus (none) (partial coverage) | Qwen | 47.6 | 2 of 4 | $0.32 / $1.28 | — | 39.2 | 51.1 | — |
| 327 | GPT-5.4 Mini (medium) (partial coverage) | OpenAI | 47.6 | 1 of 4 | $0.75 / $4.50 | — | 42.7 | — | — |
| 328 | Qwen3 Coder 480B A35b Instruct (partial coverage) | Qwen | 47.5 | 3 of 4 | — | 45.7 | 47.9 | 43.9 | — |
| 329 | Mimo v2 Flash (partial coverage) | Xiaomi | 47.5 | 2 of 4 | — | — | 45.8 | 44.2 | — |
| 330 | Gemini 3.5 Flash Lite (minimal) (partial coverage) | Google | 47.5 | 1 of 4 | $0.30 / $2.50 | — | — | 42.4 | — |
| 331 | o4 Mini (partial coverage) | OpenAI | 47.5 | 4 of 4 | $1.10 / $4.40 | 41.6 | 49.8 | 51.1 | 42.5 |
| 332 | KAT-Coder-Pro V1 (partial coverage) | Kwaipilot | 47.5 | 1 of 4 | — | — | 42.4 | — | — |
| 333 | DeepSeek V4 Flash 0731 (high) (partial coverage) | DeepSeek | 47.5 | 1 of 4 | $0.14 / $0.28 | — | 42.2 | — | — |
| 334 | O1 Mini (high) | OpenAI | 47.5 | 1 of 4 | — | — | — | 42.2 | — |
| 335 | Qwen 2.5 (max) (partial coverage) | Qwen | 47.5 | 2 of 4 | — | — | 47.3 | 42.3 | — |
| 336 | Qwen VL (max) | Qwen | 47.4 | 1 of 4 | — | — | — | — | 42.0 |
| 337 | Qwen3 235B A22B (partial coverage) | Qwen | 47.4 | 2 of 4 | $0.455 / $1.82 | — | 52.5 | 36.8 | — |
| 338 | Qwen3.5-35B-A3B (partial coverage) | Qwen | 47.4 | 2 of 4 | $0.225 / $1.80 | — | 41.5 | 47.8 | — |
| 339 | Claude 3.7 Sonnet (thinking 32K) (partial coverage) | Anthropic | 47.4 | 3 of 4 | — | — | 54.3 | 45.2 | 37.1 |
| 340 | Gemini 3.1 Flash Lite (high) (partial coverage) | Google | 47.4 | 2 of 4 | $0.25 / $1.50 | 38.0 | — | 51.1 | — |
| 341 | Ling Flash 2.0 (partial coverage) | inclusionAI | 47.4 | 2 of 4 | — | — | 48.3 | 40.9 | — |
| 342 | Claude Opus 4.6 (low) (partial coverage) | Anthropic | 47.4 | 1 of 4 | $5.00 / $25.00 | — | 41.8 | — | — |
| 343 | Gemini 3 Pro (high) (partial coverage) | Google | 47.3 | 1 of 4 | — | 41.8 | — | — | — |
| 344 | Grok 4 Fast (reasoning) (partial coverage) | xAI | 47.3 | 2 of 4 | — | — | 39.0 | 50.1 | — |
| 345 | Nemotron 3 Ultra (partial coverage) | NVIDIA | 47.3 | 3 of 4 | $0.60 / $3.60 | 20.3 | 59.3 | 56.8 | — |
| 346 | GPT-5.6 Terra (medium) (partial coverage) | OpenAI | 47.3 | 1 of 4 | $1.00 / $6.00 | — | 41.6 | — | — |
| 347 | GPT-5.4 Mini (xhigh) (partial coverage) | OpenAI | 47.3 | 3 of 4 | $0.75 / $4.50 | 36.9 | 47.9 | 51.4 | — |
| 348 | GPT-5 Nano (low) (partial coverage) | OpenAI | 47.2 | 1 of 4 | $0.05 / $0.40 | — | — | 41.4 | — |
| 349 | GPT-5 (minimal) (partial coverage) | OpenAI | 47.2 | 1 of 4 | $1.25 / $10.00 | — | — | 41.4 | — |
| 350 | GPT-5.1-Codex-Mini (partial coverage) | OpenAI | 47.2 | 1 of 4 | $0.25 / $2.00 | — | 41.5 | — | — |
| 351 | GPT-5.4 Nano (none) (partial coverage) | OpenAI | 47.2 | 1 of 4 | $0.20 / $1.25 | — | — | 41.4 | — |
| 352 | Grok 3 Mini Beta (partial coverage) | xAI | 47.2 | 2 of 4 | — | — | 45.4 | 43.0 | — |
| 353 | Claude Opus 4 (27K) | Anthropic | 47.2 | 1 of 4 | $15.00 / $75.00 | — | — | 41.3 | — |
| 354 | Grok 4 0709 (partial coverage) | xAI | 47.0 | 3 of 4 | — | — | 51.3 | 48.0 | 35.6 |
| 355 | Claude Opus 4 (partial coverage) | Anthropic | 47.0 | 4 of 4 | $15.00 / $75.00 | 55.7 | 50.9 | 38.4 | 36.7 |
| 356 | Qwen 2.5 (partial coverage) | Qwen | 47.0 | 1 of 4 | — | — | 40.6 | — | — |
| 357 | Claude Sonnet 4 (partial coverage) | Anthropic | 47.0 | 4 of 4 | $3.00 / $15.00 | 52.0 | 58.8 | 35.7 | 35.1 |
| 358 | Claude Opus 4.8 (medium) | Anthropic | 46.9 | 1 of 4 | $5.00 / $25.00 | — | 40.6 | — | — |
| 359 | Gemini 3.1 Flash Lite (low) (partial coverage) | Google | 46.9 | 1 of 4 | $0.25 / $1.50 | — | — | 40.6 | — |
| 360 | Gemini 3.1 Pro Preview (high) (partial coverage) | Google | 46.9 | 3 of 4 | $2.00 / $12.00 | 43.5 | 29.2 | 61.6 | — |
| 361 | O1 Mini (partial coverage) | OpenAI | 46.9 | 2 of 4 | — | — | 45.5 | 41.9 | — |
| 362 | QwQ 32B (partial coverage) | Qwen | 46.9 | 2 of 4 | — | — | 44.9 | 42.6 | — |
| 363 | Claude Sonnet 4.6 (low) (partial coverage) | Anthropic | 46.9 | 1 of 4 | $3.00 / $15.00 | — | 40.5 | — | — |
| 364 | Mistral Large 3 (partial coverage) | Mistral AI | 46.9 | 3 of 4 | — | — | 45.0 | 48.8 | 40.4 |
| 365 | Claude Sonnet 4.5 (thinking) (partial coverage) | Anthropic | 46.9 | 1 of 4 | $3.00 / $15.00 | 40.3 | — | — | — |
| 366 | Hunyuan TurboS (partial coverage) | Tencent | 46.8 | 2 of 4 | — | — | 46.8 | 40.3 | — |
| 367 | Kimi K2 Instruct (partial coverage) | Moonshot AI | 46.8 | 2 of 4 | — | 40.8 | 46.2 | — | — |
| 368 | GPT-5.6 Luna (none) (partial coverage) | OpenAI | 46.8 | 1 of 4 | $0.10 / $0.60 | — | — | 40.1 | — |
| 369 | MiniMax M3 (none) (partial coverage) | MiniMax | 46.8 | 1 of 4 | $0.30 / $1.20 | — | 40.1 | — | — |
| 370 | GLM 4.6 (partial coverage) | Z.ai | 46.7 | 3 of 4 | $0.55 / $2.20 | 45.7 | 44.9 | 42.9 | — |
| 371 | Ring Flash 2.0 (partial coverage) | inclusionAI | 46.7 | 2 of 4 | — | — | 46.4 | 40.1 | — |
| 372 | Nova 2 Lite (partial coverage) | Amazon | 46.7 | 2 of 4 | $0.30 / $2.50 | — | 46.9 | 39.5 | — |
| 373 | O1 Mini (medium) | OpenAI | 46.7 | 1 of 4 | — | — | — | 39.7 | — |
| 374 | Claude 3.7 Sonnet (64K) | Anthropic | 46.6 | 1 of 4 | — | — | — | 39.6 | — |
| 375 | Qwen3 30B A3B (partial coverage) | Qwen | 46.6 | 2 of 4 | $0.12 / $0.50 | — | 45.3 | 41.0 | — |
| 376 | MiniMax M2.7 (partial coverage) | MiniMax | 46.6 | 3 of 4 | $0.30 / $1.20 | 24.4 | 56.0 | 52.3 | — |
| 377 | Grok 4.1 Fast (reasoning) (partial coverage) | xAI | 46.6 | 3 of 4 | — | — | 44.8 | 51.6 | 36.4 |
| 378 | Gemini 3.1 Flash Lite (minimal) (partial coverage) | Google | 46.6 | 1 of 4 | $0.25 / $1.50 | — | — | 39.5 | — |
| 379 | GPT-5 Nano (medium) (partial coverage) | OpenAI | 46.5 | 3 of 4 | $0.05 / $0.40 | 39.4 | 43.5 | 49.3 | — |
| 380 | Mimo v2 Flash (thinking) (partial coverage) | Xiaomi | 46.5 | 2 of 4 | — | — | 41.9 | 43.8 | — |
| 381 | Claude Sonnet 5 (low) (partial coverage) | Anthropic | 46.5 | 1 of 4 | $2.00 / $10.00 | — | 39.2 | — | — |
| 382 | GPT-5 Nano (minimal) (partial coverage) | OpenAI | 46.4 | 1 of 4 | $0.05 / $0.40 | — | — | 38.8 | — |
| 383 | SWE-1.6 (none) (partial coverage) | Cognition | 46.3 | 1 of 4 | — | — | 38.7 | — | — |
| 384 | Claude Sonnet 4.6 (high) (partial coverage) | Anthropic | 46.3 | 2 of 4 | $3.00 / $15.00 | — | 35.4 | 49.5 | — |
| 385 | MiniMax M2 (partial coverage) | MiniMax | 46.3 | 3 of 4 | $0.255 / $1.02 | 50.5 | 39.2 | 41.3 | — |
| 386 | Qwen2.5 VL 72B Instruct (partial coverage) | Qwen | 46.3 | 1 of 4 | — | — | — | — | 38.5 |
| 387 | Grok Code Fast 1 (partial coverage) | xAI | 46.2 | 1 of 4 | — | — | 38.5 | — | — |
| 388 | GPT-5.4 Mini (low) (partial coverage) | OpenAI | 46.2 | 1 of 4 | $0.75 / $4.50 | — | 38.3 | — | — |
| 389 | Command A (03-2025) (partial coverage) | Cohere | 46.1 | 2 of 4 | — | — | 46.0 | 38.2 | — |
| 390 | GLM 5.2 (none) (partial coverage) | Z.ai | 46.1 | 2 of 4 | $0.462 / $1.452 | — | 46.2 | 37.8 | — |
| 391 | Qwen2.5 VL 32B Instruct (partial coverage) | Qwen | 46.1 | 1 of 4 | — | — | — | — | 37.9 |
| 392 | Qwen3.5-Flash (partial coverage) | Qwen | 46.0 | 2 of 4 | $0.065 / $0.26 | — | 40.9 | 43.0 | — |
| 393 | Devstral Medium 2507 (partial coverage) | Mistral AI | 46.0 | 1 of 4 | — | — | 37.9 | — | — |
| 394 | Mistral Medium 3.5 (none) (partial coverage) | Mistral | 46.0 | 1 of 4 | $1.50 / $7.50 | — | 37.9 | — | — |
| 395 | GPT-4o (partial coverage) | OpenAI | 46.0 | 1 of 4 | $2.50 / $10.00 | — | 37.9 | — | — |
| 396 | Nemotron 3 Nano 30B A3B (partial coverage) | NVIDIA | 46.0 | 2 of 4 | $0.05 / $0.20 | — | 43.2 | 40.6 | — |
| 397 | Trinity Large Thinking (partial coverage) | Arcee AI | 46.0 | 2 of 4 | $0.22 / $0.85 | — | 38.7 | 45.1 | — |
| 398 | GPT-5 Mini (partial coverage) | OpenAI | 46.0 | 2 of 4 | $0.25 / $2.00 | 46.8 | 36.9 | — | — |
| 399 | gpt-oss-20b (partial coverage) | OpenAI | 46.0 | 2 of 4 | $0.03 / $0.13 | — | 44.0 | 39.6 | — |
| 400 | Claude Sonnet 4 (59K) (partial coverage) | Anthropic | 46.0 | 1 of 4 | $3.00 / $15.00 | — | — | 37.6 | — |
| 401 | DeepSeek V3 0324 (partial coverage) | DeepSeek | 45.9 | 2 of 4 | $0.27 / $1.12 | — | 49.8 | 33.6 | — |
| 402 | GPT-5.4 Mini (none) (partial coverage) | OpenAI | 45.9 | 1 of 4 | $0.75 / $4.50 | — | — | 37.5 | — |
| 403 | Hunyuan Vision 1.5 (thinking) (partial coverage) | Tencent | 45.9 | 2 of 4 | — | — | 52.0 | — | 31.3 |
| 404 | DeepSeek V4 Pro (none) (partial coverage) | DeepSeek | 45.8 | 2 of 4 | $1.168 / $2.336 | — | 41.4 | 41.4 | — |
| 405 | Grok 4.3 (high) | SpaceXAI | 45.7 | 2 of 4 | $1.25 / $2.50 | 30.9 | — | 51.8 | — |
| 406 | OLMo 3.1 32B Instruct (partial coverage) | Allen Institute for AI | 45.6 | 2 of 4 | — | — | 44.4 | 37.6 | — |
| 407 | OLMo 3 32B Think (partial coverage) | Allen Institute for AI | 45.6 | 2 of 4 | — | — | 43.5 | 38.5 | — |
| 408 | Magistral Medium 2506 (partial coverage) | Mistral AI | 45.5 | 2 of 4 | — | — | 45.7 | 36.0 | — |
| 409 | Gemini 2.5 Flash (partial coverage) | Google | 45.5 | 4 of 4 | $0.30 / $2.50 | 38.6 | 41.4 | 44.1 | 48.4 |
| 410 | GPT-5.6 Luna (medium) (partial coverage) | OpenAI | 45.3 | 1 of 4 | $0.10 / $0.60 | — | 35.7 | — | — |
| 411 | Devstral 2 (partial coverage) | Mistral AI | 45.3 | 2 of 4 | — | 43.8 | 37.2 | — | — |
| 412 | Mercury 2 (partial coverage) | Inception | 45.3 | 1 of 4 | $0.25 / $0.75 | — | 35.6 | — | — |
| 413 | GLM 4.5 (partial coverage) | Z.ai | 45.3 | 3 of 4 | $0.60 / $2.20 | 44.5 | 48.8 | 32.7 | — |
| 414 | Claude 3.7 Sonnet (partial coverage) | Anthropic | 45.1 | 4 of 4 | — | 42.3 | 60.3 | 33.9 | 33.9 |
| 415 | Mistral Medium 2508 (partial coverage) | Mistral AI | 45.1 | 3 of 4 | — | — | 54.7 | 47.3 | 23.0 |
| 416 | Claude Opus 4.1 (16K) (partial coverage) | Anthropic | 45.1 | 1 of 4 | $15.00 / $75.00 | — | — | 34.9 | — |
| 417 | o3 Pro (partial coverage) | OpenAI | 45.0 | 1 of 4 | $20.00 / $80.00 | 34.6 | — | — | — |
| 418 | GPT-5.1 (none) (partial coverage) | OpenAI | 44.9 | 1 of 4 | $1.25 / $10.00 | — | — | 34.5 | — |
| 419 | Gemma 3 12B (partial coverage) | Google | 44.8 | 2 of 4 | $0.05 / $0.15 | — | 39.9 | 38.9 | — |
| 420 | Gemini 1.5 Pro 002 (partial coverage) | Google | 44.7 | 3 of 4 | — | — | 42.5 | 35.7 | 44.9 |
| 421 | GPT-4.1 (partial coverage) | OpenAI | 44.6 | 4 of 4 | $2.00 / $8.00 | 40.1 | 47.2 | 35.9 | 44.2 |
| 422 | GLM 5.1 (max) (partial coverage) | Z.ai | 44.6 | 1 of 4 | $0.966 / $3.036 | 33.4 | — | — | — |
| 423 | Claude Opus 4.1 (32K) (partial coverage) | Anthropic | 44.5 | 1 of 4 | $15.00 / $75.00 | — | — | 33.2 | — |
| 424 | Mistral Medium 3.5 (partial coverage) | Mistral | 44.5 | 4 of 4 | $1.50 / $7.50 | 23.8 | 48.1 | 54.8 | 39.8 |
| 425 | Gemini 3.5 Flash Lite (high) | Google | 44.4 | 1 of 4 | $0.30 / $2.50 | — | — | 32.9 | — |
| 426 | OLMo 3.1 32B Think (partial coverage) | Allen Institute for AI | 44.3 | 2 of 4 | — | — | 40.3 | 36.8 | — |
| 427 | Mistral Nemo (partial coverage) | Mistral | 44.3 | 1 of 4 | $0.019 / $0.03 | — | — | 32.8 | — |
| 428 | Gemini 2.5 Flash Lite (thinking) (partial coverage) | Google | 44.3 | 3 of 4 | $0.10 / $0.40 | — | 44.6 | 42.8 | 33.9 |
| 429 | Qwen (max) | Qwen | 44.3 | 1 of 4 | — | — | — | 32.6 | — |
| 430 | GPT-5.6 Luna (low) (partial coverage) | OpenAI | 44.3 | 2 of 4 | $0.10 / $0.60 | — | 30.7 | 46.2 | — |
| 431 | GLM 4.6V (partial coverage) | Z.ai | 44.3 | 2 of 4 | $0.30 / $0.90 | — | 49.0 | — | 27.8 |
| 432 | Gemini 2.5 Flash Lite (nothinking) (partial coverage) | Google | 44.2 | 3 of 4 | $0.10 / $0.40 | — | 47.0 | 42.5 | 31.2 |
| 433 | gpt-oss-120b (partial coverage) | OpenAI | 44.1 | 3 of 4 | $0.03 / $0.17 | 37.9 | 37.8 | 44.6 | — |
| 434 | Gemini 2.0 Flash Lite Preview 02 05 (partial coverage) | Google | 44.0 | 3 of 4 | — | — | 40.7 | 39.3 | 39.6 |
| 435 | Grok 2 1212 | xAI | 43.9 | 1 of 4 | — | — | — | 31.5 | — |
| 436 | Claude Haiku 4.5 (partial coverage) | Anthropic | 43.9 | 3 of 4 | $1.00 / $5.00 | 33.4 | 43.4 | 42.3 | — |
| 437 | DeepSeek V3 | DeepSeek | 43.7 | 2 of 4 | $0.2574 / $1.0287 | — | 41.1 | 33.2 | — |
| 438 | Solar Pro 4 (partial coverage) | Upstage | 43.6 | 2 of 4 | $0.03 / $0.12 | 24.9 | 49.3 | — | — |
| 439 | Gemma 3n E4B (partial coverage) | Google | 43.6 | 2 of 4 | — | — | 39.4 | 34.6 | — |
| 440 | GPT-4.1 Mini (partial coverage) | OpenAI | 43.5 | 4 of 4 | $0.40 / $1.60 | 37.1 | 41.1 | 42.4 | 40.2 |
| 441 | Qwen2.5 72B Instruct (partial coverage) | Qwen | 43.4 | 2 of 4 | $0.36 / $0.40 | — | 42.3 | 31.1 | — |
| 442 | Claude Opus 4.7 (medium) | Anthropic | 43.3 | 1 of 4 | $5.00 / $25.00 | — | 29.6 | — | — |
| 443 | Gemma 3 4B (partial coverage) | Google | 43.2 | 2 of 4 | — | — | 38.3 | 34.3 | — |
| 444 | Magistral Small 2506 (partial coverage) | Mistral AI | 43.2 | 1 of 4 | — | — | — | 29.3 | — |
| 445 | Command R+ (08-2024) (partial coverage) | Cohere | 43.2 | 2 of 4 | $2.50 / $10.00 | — | 38.7 | 33.8 | — |
| 446 | GLM 5.2 (high) | Z.ai | 43.1 | 1 of 4 | $0.462 / $1.452 | — | 29.2 | — | — |
| 447 | SWE Llama | Princeton NLP | 43.1 | 1 of 4 | — | — | 29.0 | — | — |
| 448 | Amazon Nova Micro v1.0 (partial coverage) | Amazon | 43.1 | 2 of 4 | — | — | 39.0 | 33.2 | — |
| 449 | GLM 4.5V (partial coverage) | Z.ai | 43.1 | 3 of 4 | $0.60 / $1.80 | — | 47.6 | 41.6 | 25.9 |
| 450 | Grok Build 0.1 | SpaceXAI | 43.1 | 1 of 4 | $1.00 / $2.00 | 28.9 | — | — | — |
| 451 | OLMo 2 0325 32B Instruct (partial coverage) | Allen Institute for AI | 43.0 | 2 of 4 | — | — | 38.5 | 33.3 | — |
| 452 | GPT-5 Nano (high) (partial coverage) | OpenAI | 43.0 | 3 of 4 | $0.05 / $0.40 | — | 44.7 | 43.9 | 26.1 |
| 453 | Command R (08-2024) (partial coverage) | Cohere | 43.0 | 2 of 4 | $0.15 / $0.60 | — | 38.8 | 32.9 | — |
| 454 | GPT-4o (2024-05-13) (partial coverage) | OpenAI | 43.0 | 3 of 4 | $5.00 / $15.00 | — | 41.8 | 29.2 | 43.5 |
| 455 | Step 3 (partial coverage) | StepFun | 43.0 | 3 of 4 | — | — | 48.0 | 42.0 | 24.4 |
| 456 | Granite 4.1 8B (partial coverage) | IBM | 42.9 | 2 of 4 | $0.05 / $0.10 | — | 32.3 | 39.0 | — |
| 457 | QwQ 32B Preview (partial coverage) | Qwen | 42.8 | 2 of 4 | — | — | 37.9 | 33.0 | — |
| 458 | GPT 4 1106 Preview (partial coverage) | OpenAI | 42.6 | 2 of 4 | — | — | 39.1 | 31.3 | — |
| 459 | SWE Llama 13B | Princeton NLP | 42.5 | 1 of 4 | — | — | 27.4 | — | — |
| 460 | Mistral Large 2407 (partial coverage) | Mistral | 42.4 | 2 of 4 | $2.00 / $6.00 | — | 42.0 | 27.4 | — |
| 461 | Amazon Nova Pro v1.0 (partial coverage) | Amazon | 42.3 | 3 of 4 | — | — | 40.9 | 34.9 | 35.4 |
| 462 | GPT-4.1 Nano (partial coverage) | OpenAI | 42.2 | 3 of 4 | $0.10 / $0.40 | — | 44.3 | 29.8 | 36.5 |
| 463 | Llama 3.3 70B Instruct (partial coverage) | Meta | 41.8 | 2 of 4 | $0.10 / $0.32 | — | 41.0 | 25.9 | — |
| 464 | Amazon Nova Lite v1.0 (partial coverage) | Amazon | 41.7 | 3 of 4 | — | — | 39.2 | 34.0 | 35.1 |
| 465 | Mistral Large 2411 (partial coverage) | Mistral AI | 41.6 | 2 of 4 | — | — | 41.1 | 25.0 | — |
| 466 | Mistral Small 2506 (partial coverage) | Mistral AI | 41.5 | 3 of 4 | — | — | 48.4 | 39.8 | 19.0 |
| 467 | GPT-4o (2024-11-20) (partial coverage) | OpenAI | 41.4 | 3 of 4 | $2.50 / $10.00 | 36.4 | 43.2 | 27.2 | — |
| 468 | Mixtral 8x22B Instruct (partial coverage) | Mistral | 41.4 | 2 of 4 | $2.00 / $6.00 | — | 38.4 | 26.9 | — |
| 469 | Mistral Small 24B Instruct 2501 (partial coverage) | Mistral AI | 41.3 | 2 of 4 | — | — | 39.6 | 25.3 | — |
| 470 | Gemini 1.5 Pro 001 (partial coverage) | Google | 41.3 | 3 of 4 | — | — | 41.3 | 26.7 | 38.2 |
| 471 | Gemini 1.5 Flash 002 (partial coverage) | Google | 41.3 | 3 of 4 | — | — | 39.8 | 26.1 | 40.2 |
| 472 | GPT 3.5 | OpenAI | 41.0 | 1 of 4 | — | — | 22.7 | — | — |
| 473 | Claude 3.5 Sonnet | Anthropic | 41.0 | 3 of 4 | — | — | 48.8 | 26.7 | 29.2 |
| 474 | Llama 3.1 70B Instruct (partial coverage) | Meta | 41.0 | 2 of 4 | $0.40 / $0.40 | — | 40.2 | 23.4 | — |
| 475 | Mistral Medium 2505 (partial coverage) | Mistral AI | 40.9 | 3 of 4 | — | — | 50.9 | 33.6 | 20.0 |
| 476 | Hunyuan Large Vision (partial coverage) | Tencent | 40.9 | 3 of 4 | — | — | 42.4 | 35.8 | 26.0 |
| 477 | Molmo 2 8B | Allen Institute for AI | 40.8 | 1 of 4 | — | — | — | — | 22.3 |
| 478 | Step 1o Turbo 202506 (partial coverage) | StepFun | 40.6 | 3 of 4 | — | — | 41.7 | 37.2 | 23.9 |
| 479 | Gemma 3 27B (partial coverage) | Google | 40.5 | 3 of 4 | $0.08 / $0.45 | — | 42.7 | 35.8 | 24.0 |
| 480 | GPT-4o (2024-08-06) (partial coverage) | OpenAI | 40.2 | 3 of 4 | $2.50 / $10.00 | — | 36.9 | 26.4 | 37.6 |
| 481 | Gemini 1.5 Flash 8B 001 (partial coverage) | Google | 40.1 | 3 of 4 | — | — | 38.1 | 26.0 | 36.3 |
| 482 | Qwen2.5 Coder 32B Instruct (partial coverage) | Qwen | 40.1 | 3 of 4 | — | 33.4 | 31.6 | 35.2 | — |
| 483 | GPT-4 Turbo (partial coverage) | OpenAI | 39.9 | 3 of 4 | $10.00 / $30.00 | — | 34.1 | 27.9 | 37.4 |
| 484 | Llama 3.1 8B Instruct (partial coverage) | Meta | 39.8 | 2 of 4 | $0.05 / $0.08 | — | 38.0 | 21.0 | — |
| 485 | Claude 2 (partial coverage) | Anthropic | 39.8 | 2 of 4 | — | — | 33.7 | 25.1 | — |
| 486 | GPT-4o-mini (2024-07-18) (partial coverage) | OpenAI | 39.8 | 3 of 4 | $0.15 / $0.60 | — | 33.2 | 28.5 | 36.8 |
| 487 | Gemini 1.5 Flash 001 (partial coverage) | Google | 39.7 | 3 of 4 | — | — | 39.5 | 22.8 | 35.7 |
| 488 | Claude 3.5 Haiku (partial coverage) | Anthropic | 39.5 | 3 of 4 | — | — | 53.5 | 23.2 | 20.7 |
| 489 | Claude 3 Opus (partial coverage) | Anthropic | 39.5 | 3 of 4 | — | — | 35.0 | 26.2 | 36.0 |
| 490 | Claude 3 Sonnet (partial coverage) | Anthropic | 39.3 | 3 of 4 | — | — | 40.1 | 21.5 | 34.8 |
| 491 | Gemini 2.0 Flash 001 (partial coverage) | Google | 38.8 | 4 of 4 | — | 34.9 | 34.4 | 35.7 | 27.7 |
| 492 | Llama 4 Maverick 17B 128e Instruct (partial coverage) | Meta | 37.7 | 4 of 4 | — | 35.7 | 35.4 | 31.1 | 23.6 |
| 493 | Mistral Small 3.1 24B Instruct 2503 (partial coverage) | Mistral AI | 37.6 | 3 of 4 | — | — | 42.9 | 26.9 | 17.9 |
| 494 | Claude 3 Haiku (partial coverage) | Anthropic | 36.7 | 3 of 4 | $0.25 / $1.25 | — | 27.8 | 20.9 | 34.6 |
| 495 | Llama 4 Scout 17B 16e Instruct (partial coverage) | Meta | 36.0 | 4 of 4 | — | 34.2 | 33.7 | 26.5 | 21.1 |

## How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model ranks here on however many of the 4 text categories it has been scored in, and the coverage column says how many that is. Each category score is itself a shrunk mean of that model's benchmark percentiles, so a composite standing on one category sits near the middle of the board rather than at the top of it.

## Data sources

- [DeepSWE v1.1](https://deepswe.datacurve.ai/): Pass@1 across all 113 DeepSWE v1.1 tasks, run by Datacurve with mini-swe-agent through Pier; reasoning-effort configurations remain separate.
- [Epoch AI Benchmarking Hub](https://epoch.ai/benchmarks): CC BY 4.0. Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- [FrontierCode 1.1](https://cognition.com/frontiercode): Main weighted rubric Scores on FrontierCode 1.1's 100 hardest tasks, run by Cognition across five trials; results identify model, harness, and effort.
- [FrontierSWE](https://www.frontierswe.com/): Mean@5 Dominance across FrontierSWE's complete 17-task cohort; results identify the evaluated model and harness.
- [LiveCodeBench](https://livecodebench.github.io/leaderboard): MIT. Maintainer-published Code Generation pass@1 over the 454-problem release_v6 window (2024-08-01 through 2025-05-01), from the stable 2025-08-01 snapshot; it does not cover newer frontier families.
- [LMArena](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset): CC BY 4.0. Arena ratings by LMArena, from the public leaderboard dataset.
- [MCP Atlas](https://labs.scale.com/leaderboard/mcp_atlas): MIT. All-1,000-task Pass Rate from Scale's current Performance Comparison, using the standardized MCP loop and 100-tool-call budget.
- [Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard): Word error rates from the Hugging Face Open ASR Leaderboard.
- [SWE-bench](https://www.swebench.com): Resolve rates published by the SWE-bench maintainers.
- [Terminal-Bench 2.1](https://www.tbench.ai/leaderboard/terminal-bench/2.1): Apache-2.0. Published Accuracy over 89 Terminal-Bench 2.1 tasks, run and verified by the benchmark team; results are model plus agent harness plus effort, not model-only evaluations.
- [TTS Arena V2](https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2): Elo ratings from TTS Arena V2.
- [Warden](https://warden.sentry.dev/benchmarking): Security review results published by Warden.
