Best LLM for math
Research-grade problem sets plus competition math. Contamination-resistant private sets carry the same weight as public ones.
Aldena runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | pricein / out | benchmarks | % | % | % | % | % | elo |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 (max)anthropic/claude-fable-5:max | Anthropic | 82.0 | $10.00 / $50.00 | 4 of 6 benchmarks | 100.0% | 87.0% | 100.0% | 99.7% | ||
| 2 | Claude Opus 5 (max)anthropic/claude-opus-5:max | Anthropic | 80.5 | $5.00 / $25.00 | 4 of 6 benchmarks | 85.6% | 73.2% | 98.9% | 1556 | ||
| 3 | GPT-5.6 Sol (max)openai/gpt-5.6-sol:max | OpenAI | 79.3 | $5.00 / $30.00 | 3 of 6 benchmarks | 89.1% | 82.9% | 100.0% | |||
| 4 | GPT-5.4 (high)openai/gpt-5.4:high | OpenAI | 78.8 | $2.50 / $15.00 | 4 of 6 benchmarks | 80.0% | 50.0% | 97.8% | 1494 | ||
| 5 | GPT-5.6 Terra (max)openai/gpt-5.6-terra:max | OpenAI | 77.0 | $1.00 / $6.00 | 3 of 6 benchmarks | 86.0% | 70.7% | 99.7% | |||
| 6 | Claude Opus 4.8 (max)anthropic/claude-opus-4.8:max | Anthropic | 74.8 | $5.00 / $25.00 | 4 of 6 benchmarks | 47.2% | 80.0% | 56.1% | 98.3% | ||
| 7 | GPT-5.6 Luna (max)openai/gpt-5.6-luna:max | OpenAI | 74.6 | $0.10 / $0.60 | 3 of 6 benchmarks | 82.1% | 61.0% | 98.3% | |||
| 8 | GPT 5.5 Pro Pre Release (xhigh)openai/gpt-5.5-pro-pre-release:xhigh | OpenAI | 74.2 | — | 3 of 6 benchmarks | 51.0% | 39.6% | 100.0% | |||
| 9 | Kimi K3 (max)moonshotai/kimi-k3:max | MoonshotAI | 74.1 | $3.00 / $15.00 | 4 of 6 benchmarks | 72.2% | 39.0% | 97.2% | 1495 | ||
| 10 | GPT 5.5 Pre Release (xhigh)openai/gpt-5.5-pre-release:xhigh | OpenAI | 73.7 | — | 3 of 6 benchmarks | 51.7% | 35.4% | 100.0% | |||
| 11 | GPT-5.5 Pro (xhigh)openai/gpt-5.5-pro:xhigh | OpenAI | 73.5 | $30.00 / $180.00 | 2 of 6 benchmarks | 87.7% | 78.0% | ||||
| 12 | Gemini 3.5 Flash (high)google/gemini-3.5-flash:high | 72.8 | $1.50 / $9.00 | 5 of 6 benchmarks | 80.0% | 62.8% | 26.8% | 95.6% | 1507 | ||
| 13 | Qwen3.8 Max (xhigh)qwen/qwen3.8-max:xhigh | Qwen | 72.7 | $2.00 / $6.00 | 3 of 6 benchmarks | 74.7% | 46.3% | 99.4% | |||
| 14 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 72.6 | $2.00 / $12.00 | 5 of 6 benchmarks | 88.9% | 59.6% | 26.8% | 95.6% | 1491 | ||
| 15 | GPT-5.4 Pro (xhigh)openai/gpt-5.4-pro:xhigh | OpenAI | 72.5 | $30.00 / $180.00 | 3 of 6 benchmarks | 50.0% | 82.5% | 58.5% | |||
| 16 | GPT-5.4 (xhigh)openai/gpt-5.4:xhigh | OpenAI | 72.5 | $2.50 / $15.00 | 4 of 6 benchmarks | 47.6% | 78.6% | 49.0% | 95.3% | ||
| 17 | GPT-5.5 (xhigh)openai/gpt-5.5:xhigh | OpenAI | 70.9 | $5.00 / $30.00 | 2 of 6 benchmarks | 85.3% | 72.5% | ||||
| 18 | GPT-5.2 (high)openai/gpt-5.2:high | OpenAI | 70.6 | $1.75 / $14.00 | 4 of 6 benchmarks | 60.0% | 18.8% | 96.1% | 1458 | ||
| 19 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen | 69.9 | $1.475 / $4.425 | 4 of 6 benchmarks | 64.6% | 34.1% | 95.6% | 1489 | ||
| 20 | Claude Opus 4.6 (64K)anthropic/claude-opus-4.6:64k | Anthropic | 69.6 | $5.00 / $25.00 | 3 of 6 benchmarks | 90.0% | 20.8% | 94.4% | |||
| 21 | Gemini 3.7 Flash (high)google/gemini-3.7-flash:high | 69.2 | $0.375 / $1.875 | 3 of 6 benchmarks | 71.6% | 36.6% | 97.2% | ||||
| 22 | GPT-5.2 (xhigh)openai/gpt-5.2:xhigh | OpenAI | 68.8 | $1.75 / $14.00 | 4 of 6 benchmarks | 40.7% | 67.4% | 31.7% | 96.1% | ||
| 23 | Claude Opus 4.6 (max)anthropic/claude-opus-4.6:max | Anthropic | 68.7 | $5.00 / $25.00 | 4 of 6 benchmarks | 90.0% | 66.0% | 26.8% | 91.1% | ||
| 24 | GPT 5.5 Pro Pre Release (high)openai/gpt-5.5-pro-pre-release:high | OpenAI | 68.4 | — | 2 of 6 benchmarks | 52.4% | 39.6% | ||||
| 25 | Grok 4.6 (xhigh)x-ai/grok-4.6:xhigh | SpaceXAI | 68.1 | $2.00 / $6.00 | 3 of 6 benchmarks | 66.0% | 31.7% | 99.2% | |||
| 26 | Claude Opus 4.6 (32K)anthropic/claude-opus-4.6:32k | Anthropic | 67.9 | $5.00 / $25.00 | 3 of 6 benchmarks | 80.0% | 20.8% | 93.1% | |||
| 27 | Kimi K2.6moonshotai/kimi-k2.6 | MoonshotAI | 67.2 | $0.5415 / $2.28 | 5 of 6 benchmarks | 39.0% | 57.2% | 25.6% | 96.1% | 1478 | |
| 28 | DeepSeek V4 Pro (high)deepseek/deepseek-v4-pro:high | DeepSeek | 67.0 | $1.168 / $2.336 | 2 of 6 benchmarks | 95.6% | 1469 | ||||
| 29 | Gemini 3.6 Flash (high)google/gemini-3.6-flash:high | 66.6 | $0.75 / $3.75 | 4 of 6 benchmarks | 59.0% | 21.9% | 94.2% | 1513 | |||
| 30 | GPT-5.2 (medium)openai/gpt-5.2:medium | OpenAI | 66.3 | $1.75 / $14.00 | 3 of 6 benchmarks | 60.0% | 16.7% | 93.9% | |||
| 31 | GLM 5.1z-ai/glm-5.1 | Z.ai | 66.2 | $0.966 / $3.036 | 4 of 6 benchmarks | 33.5% | 12.5% | 93.3% | 1481 | ||
| 32 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 65.9 | $10.00 / $50.00 | 1 of 6 benchmarks | 1527 | |||||
| 33 | Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | 65.9 | $0.32 / $1.28 | 2 of 6 benchmarks | 93.3% | 1471 | ||||
| 34 | Claude Fable 5 (high)anthropic/claude-fable-5:high | Anthropic | 65.8 | $10.00 / $50.00 | 1 of 6 benchmarks | 100.0% | |||||
| 35 | Claude Opus 5 (high)anthropic/claude-opus-5:high | Anthropic | 65.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 1525 | |||||
| 36 | GPT-5.2 Pro (xhigh)openai/gpt-5.2-pro:xhigh | OpenAI | 65.7 | $21.00 / $168.00 | 2 of 6 benchmarks | 74.0% | 46.0% | ||||
| 37 | Claude Opus 4.6 (high)anthropic/claude-opus-4.6:high | Anthropic | 65.6 | $5.00 / $25.00 | 1 of 6 benchmarks | 1516 | |||||
| 38 | Qwen3.8 Maxqwen/qwen3.8-max | Qwen | 65.5 | $2.00 / $6.00 | 1 of 6 benchmarks | 1513 | |||||
| 39 | GPT-5 (high)openai/gpt-5:high | OpenAI | 65.5 | $1.25 / $10.00 | 6 of 6 benchmarks | 32.4% | 55.4% | 21.9% | 91.4% | 98.1% | 1435 |
| 40 | Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | 65.3 | $5.00 / $25.00 | 3 of 6 benchmarks | 38.3% | 14.6% | 1506 | |||
| 41 | Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | 65.2 | $0.325 / $1.95 | 4 of 6 benchmarks | 50.0% | 8.3% | 93.3% | 1454 | ||
| 42 | Claude Opus 4.7 (high)anthropic/claude-opus-4.7:high | Anthropic | 64.9 | $5.00 / $25.00 | 1 of 6 benchmarks | 1503 | |||||
| 43 | GPT-5.5openai/gpt-5.5 | OpenAI | 64.8 | $5.00 / $30.00 | 1 of 6 benchmarks | 1499 | |||||
| 44 | Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 64.6 | $0.50 / $3.00 | 5 of 6 benchmarks | 60.0% | 51.2% | 17.1% | 92.8% | 1476 | ||
| 45 | Muse Sparkmeta/muse-spark | Meta | 64.4 | — | 4 of 6 benchmarks | 39.0% | 14.6% | 88.9% | 1461 | ||
| 46 | Claude Opus 4.8 (high)anthropic/claude-opus-4.8:high | Anthropic | 64.4 | $5.00 / $25.00 | 1 of 6 benchmarks | 1493 | |||||
| 47 | GPT-5.5 (high)openai/gpt-5.5:high | OpenAI | 64.2 | $5.00 / $30.00 | 1 of 6 benchmarks | 1491 | |||||
| 48 | Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | 64.1 | $5.00 / $25.00 | 1 of 6 benchmarks | 1491 | |||||
| 49 | Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | 64.1 | $1.027 / $6.162 | 4 of 6 benchmarks | 50.0% | 4.2% | 91.1% | 1474 | ||
| 50 | Claude Fable 5 (low)anthropic/claude-fable-5:low | Anthropic | 63.8 | $10.00 / $50.00 | 1 of 6 benchmarks | 97.8% | |||||
| 51 | Claude Opus 4.8 (low)anthropic/claude-opus-4.8:low | Anthropic | 63.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 97.8% | |||||
| 52 | Claude Opus 5anthropic/claude-opus-5 | Anthropic | 63.8 | $5.00 / $25.00 | 1 of 6 benchmarks | 97.8% | |||||
| 53 | Grok 4.6 (high)x-ai/grok-4.6:high | SpaceXAI | 63.8 | $2.00 / $6.00 | 1 of 6 benchmarks | 97.8% | |||||
| 54 | Muse Spark 1.1meta/muse-spark-1.1 | Meta | 63.6 | $1.25 / $4.25 | 1 of 6 benchmarks | 1488 | |||||
| 55 | GPT-5 (medium)openai/gpt-5:medium | OpenAI | 63.6 | $1.25 / $10.00 | 4 of 6 benchmarks | 27.2% | 6.3% | 87.2% | 97.9% | ||
| 56 | GLM 5.2 (max)z-ai/glm-5.2:max | Z.ai | 63.6 | $0.462 / $1.452 | 4 of 6 benchmarks | 59.2% | 29.3% | 86.4% | 1474 | ||
| 57 | GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh | OpenAI | 63.5 | $0.10 / $0.60 | 1 of 6 benchmarks | 1484 | |||||
| 58 | GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh | OpenAI | 63.4 | $1.00 / $6.00 | 1 of 6 benchmarks | 1483 | |||||
| 59 | Claude Opus 4.7 (max)anthropic/claude-opus-4.7:max | Anthropic | 63.3 | $5.00 / $25.00 | 3 of 6 benchmarks | 70.2% | 31.7% | 86.7% | |||
| 60 | Hy3tencent/hy3 | Tencent | 63.2 | $0.132 / $0.528 | 1 of 6 benchmarks | 1481 | |||||
| 61 | Inklingthinkingmachines/inkling | Thinking Machines | 62.9 | $0.95 / $4.05 | 1 of 6 benchmarks | 1480 | |||||
| 62 | Grok 4.5x-ai/grok-4.5 | SpaceXAI | 62.8 | $2.00 / $6.00 | 1 of 6 benchmarks | 1479 | |||||
| 63 | Gemini 3 Progoogle/gemini-3-pro | 62.6 | — | 1 of 6 benchmarks | 1479 | ||||||
| 64 | GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh | OpenAI | 62.5 | $5.00 / $30.00 | 1 of 6 benchmarks | 1478 | |||||
| 65 | Gemini 3 Pro Previewgoogle/gemini-3-pro-preview | 62.4 | — | 3 of 6 benchmarks | 37.6% | 18.8% | 91.4% | ||||
| 66 | GPT-5.1 (high)openai/gpt-5.1:high | OpenAI | 62.3 | $1.25 / $10.00 | 4 of 6 benchmarks | 31.0% | 12.5% | 88.6% | 1456 | ||
| 67 | Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium | 62.2 | $1.50 / $9.00 | 1 of 6 benchmarks | 1478 | ||||||
| 68 | MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | Xiaomi | 62.1 | $0.435 / $0.87 | 1 of 6 benchmarks | 1477 | |||||
| 69 | Ernie 5.1baidu/ernie-5.1 | Baidu | 61.9 | — | 1 of 6 benchmarks | 1476 | |||||
| 70 | Grok 4.5 (high)x-ai/grok-4.5:high | SpaceXAI | 61.8 | $2.00 / $6.00 | 3 of 6 benchmarks | 57.2% | 24.4% | 97.8% | |||
| 71 | Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high | 61.6 | $0.50 / $3.00 | 1 of 6 benchmarks | 95.6% | ||||||
| 72 | Gemini 3.1 Pro Preview (high)google/gemini-3.1-pro-preview:high | 61.6 | $2.00 / $12.00 | 1 of 6 benchmarks | 95.6% | ||||||
| 73 | GPT-5.4 (medium)openai/gpt-5.4:medium | OpenAI | 61.6 | $2.50 / $15.00 | 1 of 6 benchmarks | 95.6% | |||||
| 74 | GPT-5.6 Sol (low)openai/gpt-5.6-sol:low | OpenAI | 61.6 | $5.00 / $30.00 | 1 of 6 benchmarks | 95.6% | |||||
| 75 | Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Qwen | 61.4 | $0.39 / $2.34 | 2 of 6 benchmarks | 88.9% | 1449 | ||||
| 76 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | 61.4 | $5.00 / $25.00 | 1 of 6 benchmarks | 1472 | |||||
| 77 | Gemma 4 31Bgoogle/gemma-4-31b-it | 61.2 | $0.10 / $0.34 | 1 of 6 benchmarks | 1472 | ||||||
| 78 | Qwen3.5 Max Previewqwen/qwen3.5-max-preview | Qwen | 60.9 | — | 1 of 6 benchmarks | 1470 | |||||
| 79 | Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 60.9 | $1.25 / $10.00 | 1 of 6 benchmarks | 95.9% | ||||||
| 80 | Kimi K2.5 (thinking)moonshotai/kimi-k2.5:thinking | MoonshotAI | 60.8 | $0.57 / $2.85 | 1 of 6 benchmarks | 1470 | |||||
| 81 | Claude Sonnet 5 (xhigh)anthropic/claude-sonnet-5:xhigh | Anthropic | 60.7 | $2.00 / $10.00 | 1 of 6 benchmarks | 94.7% | |||||
| 82 | Claude Opus 4.5 (high 32K)anthropic/claude-opus-4.5:high-32k | Anthropic | 60.7 | $5.00 / $25.00 | 1 of 6 benchmarks | 1469 | |||||
| 83 | DeepSeek V4 Flash 0731 (max)deepseek/deepseek-v4-flash-0731:max | DeepSeek | 60.4 | $0.14 / $0.28 | 3 of 6 benchmarks | 57.5% | 24.4% | 94.4% | |||
| 84 | Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high | Anthropic | 60.4 | $2.00 / $10.00 | 1 of 6 benchmarks | 1468 | |||||
| 85 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 60.2 | $0.12 / $0.40 | 1 of 6 benchmarks | 1468 | ||||||
| 86 | Qwen3 Maxqwen/qwen3-max | Qwen | 60.1 | $0.78 / $3.90 | 3 of 6 benchmarks | 73.3% | 97.1% | 1426 | |||
| 87 | Grok 4.20 Beta 0309 (reasoning)x-ai/grok-4.20-beta-0309:reasoning | xAI | 60.1 | — | 1 of 6 benchmarks | 1467 | |||||
| 88 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 59.8 | $3.00 / $15.00 | 1 of 6 benchmarks | 1463 | |||||
| 89 | GPT-5 Pro (high)openai/gpt-5-pro:high | OpenAI | 59.6 | $15.00 / $120.00 | 3 of 6 benchmarks | 60.0% | 55.8% | 19.5% | |||
| 90 | Claude Opus 5 (low)anthropic/claude-opus-5:low | Anthropic | 59.5 | $5.00 / $25.00 | 1 of 6 benchmarks | 93.3% | |||||
| 91 | Kimi K3 (high)moonshotai/kimi-k3:high | MoonshotAI | 59.5 | $3.00 / $15.00 | 1 of 6 benchmarks | 93.3% | |||||
| 92 | GPT-5.4openai/gpt-5.4 | OpenAI | 59.4 | $2.50 / $15.00 | 1 of 6 benchmarks | 1460 | |||||
| 93 | Claude Sonnet 4.5 (high 32K)anthropic/claude-sonnet-4.5:high-32k | Anthropic | 58.9 | $3.00 / $15.00 | 1 of 6 benchmarks | 1455 | |||||
| 94 | Gemini 3 Flash Preview (thinking minimal)google/gemini-3-flash-preview:thinking-minimal | 58.6 | $0.50 / $3.00 | 1 of 6 benchmarks | 1453 | ||||||
| 95 | Claude Sonnet 5 (max)anthropic/claude-sonnet-5:max | Anthropic | 58.6 | $2.00 / $10.00 | 3 of 6 benchmarks | 65.6% | 29.3% | 80.0% | |||
| 96 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Meta | 58.4 | $0.35 / $1.50 | 1 of 6 benchmarks | 1453 | |||||
| 97 | Grok 4.20 Multi Agent Beta 0309x-ai/grok-4.20-multi-agent-beta-0309 | xAI | 58.4 | — | 1 of 6 benchmarks | 1453 | |||||
| 98 | Mimo v2 Proxiaomi/mimo-v2-pro | Xiaomi | 58.2 | — | 1 of 6 benchmarks | 1452 | |||||
| 99 | GPT-5.2 Chatopenai/gpt-5.2-chat | OpenAI | 58.1 | $1.75 / $14.00 | 1 of 6 benchmarks | 1452 | |||||
| 100 | Dola Seed 2.0 Probytedance/dola-seed-2.0-pro | ByteDance | 57.9 | — | 1 of 6 benchmarks | 1451 | |||||
| 101 | GPT-5 Mini (high)openai/gpt-5-mini:high | OpenAI | 57.9 | $0.25 / $2.00 | 6 of 6 benchmarks | 27.2% | 46.7% | 12.2% | 86.7% | 97.8% | 1405 |
| 102 | Qwen3.6 27Bqwen/qwen3.6-27b | Qwen | 57.9 | $0.60 / $3.60 | 1 of 6 benchmarks | 91.1% | |||||
| 103 | o3openai/o3 | OpenAI | 57.6 | $2.00 / $8.00 | 1 of 6 benchmarks | 1447 | |||||
| 104 | Kimi K2p5moonshotai/kimi-k2p5 | Moonshot AI | 57.6 | — | 3 of 6 benchmarks | 27.9% | 4.2% | 92.2% | |||
| 105 | GPT-5.4 Mini (high)openai/gpt-5.4-mini:high | OpenAI | 57.6 | $0.75 / $4.50 | 4 of 6 benchmarks | 50.0% | 2.1% | 87.2% | 1439 | ||
| 106 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | DeepSeek | 57.5 | $1.168 / $2.336 | 1 of 6 benchmarks | 1445 | |||||
| 107 | Claude Opus 4.1 (thinking 16K)anthropic/claude-opus-4.1:thinking-16k | Anthropic | 57.4 | $15.00 / $75.00 | 1 of 6 benchmarks | 1444 | |||||
| 108 | GLM 5V Turboz-ai/glm-5v-turbo | Z.ai | 57.2 | $1.20 / $4.00 | 1 of 6 benchmarks | 1443 | |||||
| 109 | Grok 4.1 (thinking)x-ai/grok-4.1:thinking | xAI | 56.9 | — | 1 of 6 benchmarks | 1442 | |||||
| 110 | Gemini 3.5 Flash (low)google/gemini-3.5-flash:low | 56.8 | $1.50 / $9.00 | 1 of 6 benchmarks | 88.9% | ||||||
| 111 | GPT-5.6 Terra (low)openai/gpt-5.6-terra:low | OpenAI | 56.8 | $1.00 / $6.00 | 1 of 6 benchmarks | 88.9% | |||||
| 112 | gpt-oss-120b (high)openai/gpt-oss-120b:high | OpenAI | 56.8 | $0.03 / $0.17 | 1 of 6 benchmarks | 88.9% | |||||
| 113 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | NVIDIA | 56.8 | $0.60 / $3.60 | 1 of 6 benchmarks | 1442 | |||||
| 114 | GPT-5 Mini (medium)openai/gpt-5-mini:medium | OpenAI | 56.8 | $0.25 / $2.00 | 4 of 6 benchmarks | 20.3% | 4.2% | 78.3% | 96.8% | ||
| 115 | MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | 56.6 | $0.14 / $0.28 | 1 of 6 benchmarks | 1441 | |||||
| 116 | Kimi K2 Thinking Turbomoonshotai/kimi-k2-thinking-turbo | Moonshot AI | 56.5 | — | 3 of 6 benchmarks | 20.0% | 83.1% | 1436 | |||
| 117 | DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high | DeepSeek | 56.4 | $0.0643 / $0.1285 | 1 of 6 benchmarks | 1441 | |||||
| 118 | Claude Opus 4.7 (xhigh)anthropic/claude-opus-4.7:xhigh | Anthropic | 56.3 | $5.00 / $25.00 | 3 of 6 benchmarks | 43.8% | 0.0% | 97.8% | |||
| 119 | Kimi K2.5 Instantmoonshotai/kimi-k2.5-instant | Moonshot AI | 56.1 | — | 1 of 6 benchmarks | 1440 | |||||
| 120 | Ernie 5.0 0110baidu/ernie-5.0-0110 | Baidu | 55.8 | — | 1 of 6 benchmarks | 1437 | |||||
| 121 | o1 (medium)openai/o1:medium | OpenAI | 55.7 | $15.00 / $60.00 | 2 of 6 benchmarks | 73.3% | 94.4% | ||||
| 122 | Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 55.6 | $0.25 / $1.50 | 1 of 6 benchmarks | 1437 | ||||||
| 123 | GPT-5.2 (low)openai/gpt-5.2:low | OpenAI | 55.5 | $1.75 / $14.00 | 3 of 6 benchmarks | 40.0% | 6.3% | 78.9% | |||
| 124 | Kimi K2.7 Codemoonshotai/kimi-k2.7-code | MoonshotAI | 55.3 | $0.71 / $3.50 | 3 of 6 benchmarks | 54.0% | 12.2% | 95.6% | |||
| 125 | Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | Qwen | 55.3 | $0.15 / $1.00 | 1 of 6 benchmarks | 86.7% | |||||
| 126 | Qwen3.7 Flashqwen/qwen3.7-flash | Qwen | 55.3 | $0.03 / $0.13 | 1 of 6 benchmarks | 86.7% | |||||
| 127 | Mimo v2 Omnixiaomi/mimo-v2-omni | Xiaomi | 55.2 | — | 1 of 6 benchmarks | 1435 | |||||
| 128 | GPT-5.2openai/gpt-5.2 | OpenAI | 54.9 | $1.75 / $14.00 | 1 of 6 benchmarks | 1432 | |||||
| 129 | o3 (high)openai/o3:high | OpenAI | 54.9 | $2.00 / $8.00 | 4 of 6 benchmarks | 18.7% | 2.1% | 83.9% | 97.8% | ||
| 130 | Qwen3.5 Plusqwen/qwen3.5-plus | Qwen | 54.8 | $0.30 / $1.80 | 3 of 6 benchmarks | 50.0% | 2.1% | 86.7% | |||
| 131 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | 54.8 | $1.50 / $7.50 | 1 of 6 benchmarks | 1429 | |||||
| 132 | DeepSeek V3.2 Exp (thinking)deepseek/deepseek-v3.2-exp:thinking | DeepSeek | 54.6 | $0.27 / $0.41 | 1 of 6 benchmarks | 1429 | |||||
| 133 | Claude Sonnet 4.6 (32K)anthropic/claude-sonnet-4.6:32k | Anthropic | 54.4 | $3.00 / $15.00 | 1 of 6 benchmarks | 85.8% | |||||
| 134 | o4 Mini Highopenai/o4-mini-high | OpenAI | 54.4 | $1.10 / $4.40 | 5 of 6 benchmarks | 24.8% | 36.1% | 4.9% | 81.7% | 97.8% | |
| 135 | Grok 4.1x-ai/grok-4.1 | xAI | 54.4 | — | 1 of 6 benchmarks | 1428 | |||||
| 136 | Qwen3.5-27Bqwen/qwen3.5-27b | Qwen | 54.2 | $0.195 / $1.56 | 1 of 6 benchmarks | 1428 | |||||
| 137 | GPT-5.1 (medium)openai/gpt-5.1:medium | OpenAI | 53.7 | $1.25 / $10.00 | 3 of 6 benchmarks | 26.9% | 4.2% | 85.6% | |||
| 138 | MiniMax M3minimax/minimax-m3 | MiniMax | 53.6 | $0.30 / $1.20 | 2 of 6 benchmarks | 71.1% | 1440 | ||||
| 139 | GPT-5.3 Chatopenai/gpt-5.3-chat | OpenAI | 53.6 | — | 1 of 6 benchmarks | 1427 | |||||
| 140 | Claude Opus 4.8 (none)anthropic/claude-opus-4.8:none | Anthropic | 53.5 | $5.00 / $25.00 | 1 of 6 benchmarks | 84.4% | |||||
| 141 | GPT-5.4 (low)openai/gpt-5.4:low | OpenAI | 53.5 | $2.50 / $15.00 | 1 of 6 benchmarks | 84.4% | |||||
| 142 | GPT-5.5 (low)openai/gpt-5.5:low | OpenAI | 53.5 | $5.00 / $30.00 | 1 of 6 benchmarks | 84.4% | |||||
| 143 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 53.3 | $0.0643 / $0.1285 | 1 of 6 benchmarks | 1425 | |||||
| 144 | Hunyuan Hy3 Previewtencent/hunyuan-hy3-preview | Tencent | 53.2 | — | 1 of 6 benchmarks | 1425 | |||||
| 145 | DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking | DeepSeek | 53.1 | $0.269 / $0.40 | 1 of 6 benchmarks | 1424 | |||||
| 146 | Inkling Small (xhigh)thinkingmachines/inkling-small:xhigh | Thinking Machines | 53.0 | $0.45 / $1.20 | 3 of 6 benchmarks | 46.3% | 17.1% | 90.0% | |||
| 147 | GPT-5.4 Nano (high)openai/gpt-5.4-nano:high | OpenAI | 52.9 | $0.20 / $1.25 | 5 of 6 benchmarks | 25.9% | 44.9% | 12.2% | 87.8% | 1423 | |
| 148 | Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 52.9 | $0.30 / $2.50 | 1 of 6 benchmarks | 1423 | ||||||
| 149 | o3 (medium)openai/o3:medium | OpenAI | 52.6 | $2.00 / $8.00 | 2 of 6 benchmarks | 16.9% | 84.4% | ||||
| 150 | Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Qwen | 52.6 | $0.29 / $2.40 | 1 of 6 benchmarks | 1423 | |||||
| 151 | Grok 4.20 0309 (reasoning)x-ai/grok-4.20-0309:reasoning | xAI | 52.6 | — | 3 of 6 benchmarks | 44.9% | 17.1% | 92.2% | |||
| 152 | GPT-5.1openai/gpt-5.1 | OpenAI | 52.5 | $1.25 / $10.00 | 1 of 6 benchmarks | 1423 | |||||
| 153 | o1 (high)openai/o1:high | OpenAI | 52.4 | $15.00 / $60.00 | 3 of 6 benchmarks | 9.3% | 73.3% | 94.7% | |||
| 154 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 52.3 | $0.30 / $1.20 | 1 of 6 benchmarks | 1423 | |||||
| 155 | Claude Sonnet 4.6 (medium)anthropic/claude-sonnet-4.6:medium | Anthropic | 52.3 | $3.00 / $15.00 | 1 of 6 benchmarks | 82.2% | |||||
| 156 | Gemini 3.6 Flash (low)google/gemini-3.6-flash:low | 52.3 | $0.75 / $3.75 | 1 of 6 benchmarks | 82.2% | ||||||
| 157 | Qwen3.5 397B A17B (none)qwen/qwen3.5-397b-a17b:none | Qwen | 52.3 | $0.39 / $2.34 | 1 of 6 benchmarks | 82.2% | |||||
| 158 | Claude Opus 4 (thinking 16K)anthropic/claude-opus-4:thinking-16k | Anthropic | 52.2 | $15.00 / $75.00 | 1 of 6 benchmarks | 1421 | |||||
| 159 | o3 Mini (medium)openai/o3-mini:medium | OpenAI | 52.1 | $1.10 / $4.40 | 3 of 6 benchmarks | 11.3% | 63.9% | 95.2% | |||
| 160 | Grok 4.3x-ai/grok-4.3 | SpaceXAI | 51.9 | $1.25 / $2.50 | 1 of 6 benchmarks | 1419 | |||||
| 161 | Grok 4.3 (high)x-ai/grok-4.3:high | SpaceXAI | 51.8 | $1.25 / $2.50 | 3 of 6 benchmarks | 42.8% | 14.6% | 93.3% | |||
| 162 | Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | Qwen | 51.8 | $0.09 / $0.55 | 1 of 6 benchmarks | 1418 | |||||
| 163 | Grok 4.1 Fast (reasoning)x-ai/grok-4-1-fast:reasoning | xAI | 51.6 | — | 1 of 6 benchmarks | 1418 | |||||
| 164 | Kimi K2 0905moonshotai/kimi-k2-0905 | MoonshotAI | 51.5 | $0.60 / $2.50 | 1 of 6 benchmarks | 1417 | |||||
| 165 | GPT-5.4 Mini (xhigh)openai/gpt-5.4-mini:xhigh | OpenAI | 51.4 | $0.75 / $4.50 | 3 of 6 benchmarks | 51.2% | 9.8% | 88.9% | |||
| 166 | Claude Opus 4.5 (16K)anthropic/claude-opus-4.5:16k | Anthropic | 51.4 | $5.00 / $25.00 | 3 of 6 benchmarks | 40.0% | 2.1% | 81.7% | |||
| 167 | Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | 51.4 | $5.00 / $25.00 | 4 of 6 benchmarks | 20.7% | 4.2% | 48.1% | 1465 | ||
| 168 | Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Qwen | 51.3 | $0.10 / $1.10 | 1 of 6 benchmarks | 1417 | |||||
| 169 | DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | DeepSeek | 51.2 | $0.27 / $0.41 | 1 of 6 benchmarks | 1416 | |||||
| 170 | DeepSeek V3.2deepseek/deepseek-v3.2 | DeepSeek | 51.2 | $0.269 / $0.40 | 3 of 6 benchmarks | 22.1% | 2.1% | 1429 | |||
| 171 | Gemini 3.1 Flash Lite (high)google/gemini-3.1-flash-lite:high | 51.1 | $0.25 / $1.50 | 1 of 6 benchmarks | 80.0% | ||||||
| 172 | Gemini 3.5 Flash (minimal)google/gemini-3.5-flash:minimal | 51.1 | $1.50 / $9.00 | 1 of 6 benchmarks | 80.0% | ||||||
| 173 | Gemini 3.6 Flash (minimal)google/gemini-3.6-flash:minimal | 51.1 | $0.75 / $3.75 | 1 of 6 benchmarks | 80.0% | ||||||
| 174 | Qwen3.7 Plus (none)qwen/qwen3.7-plus:none | Qwen | 51.1 | $0.32 / $1.28 | 1 of 6 benchmarks | 80.0% | |||||
| 175 | R1deepseek/deepseek-r1 | DeepSeek | 51.1 | $0.70 / $2.50 | 3 of 6 benchmarks | 53.3% | 93.0% | 1412 | |||
| 176 | o4 Miniopenai/o4-mini | OpenAI | 51.1 | $1.10 / $4.40 | 1 of 6 benchmarks | 1415 | |||||
| 177 | GLM 5z-ai/glm-5 | Z.ai | 50.9 | $0.60 / $1.92 | 4 of 6 benchmarks | 16.4% | 2.1% | 80.0% | 1443 | ||
| 178 | LongCat Flash Chatmeituan/longcat-flash-chat | Meituan | 50.9 | — | 1 of 6 benchmarks | 1415 | |||||
| 179 | DeepSeek V3.1 (thinking)deepseek/deepseek-chat-v3.1:thinking | DeepSeek | 50.8 | $0.25 / $0.95 | 1 of 6 benchmarks | 1414 | |||||
| 180 | DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepSeek | 50.6 | $0.25 / $0.95 | 1 of 6 benchmarks | 1413 | |||||
| 181 | Claude Sonnet 4.6 (16K)anthropic/claude-sonnet-4.6:16k | Anthropic | 50.6 | $3.00 / $15.00 | 2 of 6 benchmarks | 80.0% | 0.0% | ||||
| 182 | GPT 5 Chatopenai/gpt-5-chat | OpenAI | 50.3 | — | 1 of 6 benchmarks | 1413 | |||||
| 183 | DeepSeek V4 Pro (max)deepseek/deepseek-v4-pro:max | DeepSeek | 50.2 | $1.168 / $2.336 | 3 of 6 benchmarks | 45.3% | 2.4% | 96.7% | |||
| 184 | Grok 4 Fast (reasoning)x-ai/grok-4-fast:reasoning | xAI | 50.1 | — | 1 of 6 benchmarks | 1410 | |||||
| 185 | Qwen3.7 Flash (none)qwen/qwen3.7-flash:none | Qwen | 50.0 | $0.03 / $0.13 | 1 of 6 benchmarks | 77.8% | |||||
| 186 | Ernie 5.0 Preview 1203baidu/ernie-5.0-preview-1203 | Baidu | 49.9 | — | 1 of 6 benchmarks | 1409 | |||||
| 187 | Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Qwen | 49.6 | $0.26 / $1.04 | 1 of 6 benchmarks | 1408 | |||||
| 188 | Claude Sonnet 4.6 (high)anthropic/claude-sonnet-4.6:high | Anthropic | 49.5 | $3.00 / $15.00 | 1 of 6 benchmarks | 75.6% | |||||
| 189 | Step 3.5 Flashstepfun/step-3.5-flash | StepFun | 49.4 | $0.10 / $0.30 | 1 of 6 benchmarks | 1406 | |||||
| 190 | GPT-5 Nano (medium)openai/gpt-5-nano:medium | OpenAI | 49.3 | $0.05 / $0.40 | 4 of 6 benchmarks | 10.0% | 2.1% | 74.2% | 95.2% | ||
| 191 | Claude Sonnet 4.5 (59K)anthropic/claude-sonnet-4.5:59k | Anthropic | 49.1 | $3.00 / $15.00 | 2 of 6 benchmarks | 13.5% | 77.8% | ||||
| 192 | Gemma 4 31B (minimal)google/gemma-4-31b-it:minimal | 48.9 | $0.10 / $0.34 | 1 of 6 benchmarks | 73.3% | ||||||
| 193 | Mistral Large 3mistralai/mistral-large-3 | Mistral AI | 48.8 | — | 1 of 6 benchmarks | 1404 | |||||
| 194 | Chatgpt 4oopenai/chatgpt-4o | OpenAI | 48.6 | — | 1 of 6 benchmarks | 1404 | |||||
| 195 | Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Qwen | 48.5 | $0.40 / $4.00 | 1 of 6 benchmarks | 1404 | |||||
| 196 | Claude Sonnet 4 (32K)anthropic/claude-sonnet-4:32k | Anthropic | 48.1 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 197 | Claude Sonnet 4.5 (16K)anthropic/claude-sonnet-4.5:16k | Anthropic | 48.1 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 198 | Claude Sonnet 4.6 (max)anthropic/claude-sonnet-4.6:max | Anthropic | 48.1 | $3.00 / $15.00 | 1 of 6 benchmarks | 71.1% | |||||
| 199 | Claude Sonnet 4 (thinking 32K)anthropic/claude-sonnet-4:thinking-32k | Anthropic | 48.1 | $3.00 / $15.00 | 1 of 6 benchmarks | 1403 | |||||
| 200 | Grok 4 0709x-ai/grok-4-0709 | xAI | 48.0 | — | 3 of 6 benchmarks | 19.7% | 2.1% | 1427 | |||
| 201 | Hunyuan T1tencent/hunyuan-t1 | Tencent | 47.9 | — | 1 of 6 benchmarks | 1401 | |||||
| 202 | Claude Opus 4.5 (32K)anthropic/claude-opus-4.5:32k | Anthropic | 47.9 | $5.00 / $25.00 | 4 of 6 benchmarks | 20.7% | 34.4% | 4.9% | 86.1% | ||
| 203 | Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 47.9 | $1.25 / $10.00 | 2 of 6 benchmarks | 30.0% | 2.1% | |||||
| 204 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Qwen | 47.8 | $0.225 / $1.80 | 1 of 6 benchmarks | 1400 | |||||
| 205 | Gemini 2.5 Progoogle/gemini-2.5-pro | 47.7 | $1.25 / $10.00 | 6 of 6 benchmarks | 14.1% | 24.6% | 0.0% | 84.7% | 95.6% | 1441 | |
| 206 | Qwen3 32Bqwen/qwen3-32b | Qwen | 47.6 | $0.08 / $0.28 | 1 of 6 benchmarks | 1399 | |||||
| 207 | Claude Sonnet 4.5 (32K)anthropic/claude-sonnet-4.5:32k | Anthropic | 47.5 | $3.00 / $15.00 | 5 of 6 benchmarks | 15.2% | 23.9% | 2.4% | 77.8% | 97.7% | |
| 208 | Claude Haiku 4.5 (32K)anthropic/claude-haiku-4.5:32k | Anthropic | 47.4 | $1.00 / $5.00 | 4 of 6 benchmarks | 5.9% | 2.1% | 66.7% | 96.4% | ||
| 209 | Mistral Medium 2508mistralai/mistral-medium-2508 | Mistral AI | 47.3 | — | 1 of 6 benchmarks | 1398 | |||||
| 210 | Ernie 5.0 Preview 1022baidu/ernie-5.0-preview-1022 | Baidu | 47.1 | — | 1 of 6 benchmarks | 1396 | |||||
| 211 | Kimi K3 (low)moonshotai/kimi-k3:low | MoonshotAI | 47.0 | $3.00 / $15.00 | 1 of 6 benchmarks | 68.9% | |||||
| 212 | GPT-5.4 Nano (low)openai/gpt-5.4-nano:low | OpenAI | 47.0 | $0.20 / $1.25 | 1 of 6 benchmarks | 68.9% | |||||
| 213 | GPT-5.6 Sol (none)openai/gpt-5.6-sol:none | OpenAI | 47.0 | $5.00 / $30.00 | 1 of 6 benchmarks | 68.9% | |||||
| 214 | Qwen3.6 35B A3B (none)qwen/qwen3.6-35b-a3b:none | Qwen | 47.0 | $0.15 / $1.00 | 1 of 6 benchmarks | 68.9% | |||||
| 215 | MiniMax M2.5minimax/minimax-m2.5 | MiniMax | 46.9 | $0.22 / $0.90 | 1 of 6 benchmarks | 1396 | |||||
| 216 | GPT-5.1 (low)openai/gpt-5.1:low | OpenAI | 46.7 | $1.25 / $10.00 | 2 of 6 benchmarks | 17.3% | 63.9% | ||||
| 217 | DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | DeepSeek | 46.6 | $0.27 / $0.95 | 1 of 6 benchmarks | 1394 | |||||
| 218 | Qwen3 235B A22B (nothinking)qwen/qwen3-235b-a22b:nothinking | Qwen | 46.5 | $0.455 / $1.82 | 1 of 6 benchmarks | 1393 | |||||
| 219 | Inkling (xhigh)thinkingmachines/inkling:xhigh | Thinking Machines | 46.4 | $0.95 / $4.05 | 3 of 6 benchmarks | 33.3% | 4.9% | 88.9% | |||
| 220 | R1 0528deepseek/deepseek-r1-0528 | DeepSeek | 46.3 | $0.50 / $2.15 | 4 of 6 benchmarks | 0.0% | 66.4% | 96.6% | 1395 | ||
| 221 | Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Qwen | 46.2 | $0.15 / $1.20 | 1 of 6 benchmarks | 1391 | |||||
| 222 | GPT-5.6 Luna (low)openai/gpt-5.6-luna:low | OpenAI | 46.2 | $0.10 / $0.60 | 1 of 6 benchmarks | 66.7% | |||||
| 223 | Qwen3.6 27B (none)qwen/qwen3.6-27b:none | Qwen | 46.2 | $0.60 / $3.60 | 1 of 6 benchmarks | 66.7% | |||||
| 224 | MiniMax M2.1minimax/minimax-m2.1 | MiniMax | 46.1 | $0.30 / $1.20 | 1 of 6 benchmarks | 1390 | |||||
| 225 | GLM 4.5 Airz-ai/glm-4.5-air | Z.ai | 45.9 | $0.13 / $0.85 | 1 of 6 benchmarks | 1390 | |||||
| 226 | Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Qwen | 45.8 | $0.23 / $2.30 | 4 of 6 benchmarks | 20.0% | 0.0% | 86.7% | 1397 | ||
| 227 | Grok 3 Mini Beta (low)x-ai/grok-3-mini-beta:low | xAI | 45.7 | — | 3 of 6 benchmarks | 2.8% | 62.2% | 90.9% | |||
| 228 | Kimi K2 0711moonshotai/kimi-k2 | MoonshotAI | 45.6 | $0.57 / $2.30 | 1 of 6 benchmarks | 1388 | |||||
| 229 | Llama 3.3 Nemotron Super 49B v1.5nvidia/llama-3.3-nemotron-super-49b-v1.5 | NVIDIA | 45.3 | — | 1 of 6 benchmarks | 1386 | |||||
| 230 | Qwen3.6 Flashqwen/qwen3.6-flash | Qwen | 45.2 | $0.1875 / $1.125 | 3 of 6 benchmarks | 20.0% | 0.0% | 84.4% | |||
| 231 | Claude 3.7 Sonnet (thinking 32K)anthropic/claude-3-7-sonnet:thinking-32k | Anthropic | 45.2 | — | 1 of 6 benchmarks | 1385 | |||||
| 232 | Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | 45.1 | $0.22 / $0.85 | 1 of 6 benchmarks | 1384 | |||||
| 233 | GPT-5.2 (none)openai/gpt-5.2:none | OpenAI | 45.0 | $1.75 / $14.00 | 1 of 6 benchmarks | 62.2% | |||||
| 234 | INTELLECT-3prime-intellect/intellect-3 | Prime Intellect | 44.9 | — | 1 of 6 benchmarks | 1382 | |||||
| 235 | o3 Miniopenai/o3-mini | OpenAI | 44.8 | $1.10 / $4.40 | 1 of 6 benchmarks | 1382 | |||||
| 236 | o4 Mini (medium)openai/o4-mini:medium | OpenAI | 44.7 | $1.10 / $4.40 | 3 of 6 benchmarks | 19.0% | 2.1% | 73.3% | |||
| 237 | Claude 3.7 Sonnet (32K)anthropic/claude-3-7-sonnet:32k | Anthropic | 44.7 | — | 3 of 6 benchmarks | 3.5% | 53.3% | 90.0% | |||
| 238 | gpt-oss-120bopenai/gpt-oss-120b | OpenAI | 44.6 | $0.03 / $0.17 | 1 of 6 benchmarks | 1381 | |||||
| 239 | GPT 5.5 Instantopenai/gpt-5.5-instant | OpenAI | 44.6 | — | 4 of 6 benchmarks | 26.3% | 2.4% | 68.1% | 1462 | ||
| 240 | Claude Opus 4 (16K)anthropic/claude-opus-4:16k | Anthropic | 44.6 | $15.00 / $75.00 | 1 of 6 benchmarks | 60.0% | |||||
| 241 | Gemini 3.5 Flash Lite (low)google/gemini-3.5-flash-lite:low | 44.6 | $0.30 / $2.50 | 1 of 6 benchmarks | 60.0% | ||||||
| 242 | Llama 3.1 Nemotron Ultra 253B v1nvidia/llama-3.1-nemotron-ultra-253b-v1 | NVIDIA | 44.5 | — | 1 of 6 benchmarks | 1380 | |||||
| 243 | Claude Opus 4.1anthropic/claude-opus-4.1 | Anthropic | 44.3 | $15.00 / $75.00 | 3 of 6 benchmarks | 5.9% | 40.0% | 1434 | |||
| 244 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 44.3 | $0.0482 / $0.1931 | 1 of 6 benchmarks | 1380 | |||||
| 245 | Mimo v2 Flashxiaomi/mimo-v2-flash | Xiaomi | 44.2 | — | 1 of 6 benchmarks | 1377 | |||||
| 246 | Gemini 2.5 Flashgoogle/gemini-2.5-flash | 44.1 | $0.30 / $2.50 | 4 of 6 benchmarks | 4.8% | 4.2% | 70.8% | 1406 | |||
| 247 | Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | NVIDIA | 44.0 | $0.085 / $0.40 | 1 of 6 benchmarks | 1376 | |||||
| 248 | GPT-5.4 (none)openai/gpt-5.4:none | OpenAI | 44.0 | $2.50 / $15.00 | 1 of 6 benchmarks | 57.8% | |||||
| 249 | GPT-5.5 (none)openai/gpt-5.5:none | OpenAI | 44.0 | $5.00 / $30.00 | 1 of 6 benchmarks | 57.8% | |||||
| 250 | o1openai/o1 | OpenAI | 43.9 | $15.00 / $60.00 | 3 of 6 benchmarks | 31.1% | 81.7% | 1409 | |||
| 251 | GPT-5 Nano (high)openai/gpt-5-nano:high | OpenAI | 43.9 | $0.05 / $0.40 | 6 of 6 benchmarks | 20.0% | 20.0% | 2.4% | 81.1% | 94.9% | 1346 |
| 252 | Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct | Qwen | 43.9 | — | 1 of 6 benchmarks | 1376 | |||||
| 253 | o4 Mini (low)openai/o4-mini:low | OpenAI | 43.9 | $1.10 / $4.40 | 2 of 6 benchmarks | 10.7% | 57.8% | ||||
| 254 | GPT 4.5 Previewopenai/gpt-4.5-preview | OpenAI | 43.9 | — | 3 of 6 benchmarks | 37.8% | 78.6% | 1408 | |||
| 255 | Mimo v2 Flash (thinking)xiaomi/mimo-v2-flash:thinking | Xiaomi | 43.8 | — | 1 of 6 benchmarks | 1374 | |||||
| 256 | Claude Opus 4.1 (27K)anthropic/claude-opus-4.1:27k | Anthropic | 43.7 | $15.00 / $75.00 | 3 of 6 benchmarks | 7.2% | 4.2% | 68.9% | |||
| 257 | GPT-5 Mini (minimal)openai/gpt-5-mini:minimal | OpenAI | 43.6 | $0.25 / $2.00 | 1 of 6 benchmarks | 55.6% | |||||
| 258 | o3 (low)openai/o3:low | OpenAI | 43.4 | $2.00 / $8.00 | 2 of 6 benchmarks | 9.7% | 60.0% | ||||
| 259 | MiniMax M1minimax/minimax-m1 | MiniMax | 43.3 | $0.55 / $2.20 | 1 of 6 benchmarks | 1371 | |||||
| 260 | gpt-oss-20b (high)openai/gpt-oss-20b:high | OpenAI | 43.3 | $0.03 / $0.13 | 1 of 6 benchmarks | 53.9% | |||||
| 261 | Grok 3 Mini Betax-ai/grok-3-mini-beta | xAI | 43.0 | — | 1 of 6 benchmarks | 1368 | |||||
| 262 | Claude 3.7 Sonnet (16K)anthropic/claude-3-7-sonnet:16k | Anthropic | 43.0 | — | 3 of 6 benchmarks | 4.1% | 46.7% | 86.3% | |||
| 263 | Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Qwen | 43.0 | $0.065 / $0.26 | 4 of 6 benchmarks | 10.0% | 0.0% | 84.4% | 1403 | ||
| 264 | GLM 4.7 Flashz-ai/glm-4.7-flash | Z.ai | 42.9 | $0.06 / $0.40 | 1 of 6 benchmarks | 1366 | |||||
| 265 | GLM 4.6z-ai/glm-4.6 | Z.ai | 42.9 | $0.55 / $2.20 | 3 of 6 benchmarks | 3.8% | 2.1% | 1420 | |||
| 266 | Claude Sonnet 4 (16K)anthropic/claude-sonnet-4:16k | Anthropic | 42.9 | $3.00 / $15.00 | 1 of 6 benchmarks | 53.3% | |||||
| 267 | GPT-5.6 Terra (none)openai/gpt-5.6-terra:none | OpenAI | 42.9 | $1.00 / $6.00 | 1 of 6 benchmarks | 53.3% | |||||
| 268 | o1 (low)openai/o1:low | OpenAI | 42.9 | $15.00 / $60.00 | 1 of 6 benchmarks | 53.3% | |||||
| 269 | Gemini 2.5 Flash Lite (thinking)google/gemini-2.5-flash-lite:thinking | 42.8 | $0.10 / $0.40 | 1 of 6 benchmarks | 1365 | ||||||
| 270 | o3 Mini Highopenai/o3-mini-high | OpenAI | 42.8 | $1.10 / $4.40 | 6 of 6 benchmarks | 12.4% | 18.6% | 0.0% | 76.9% | 96.5% | 1405 |
| 271 | Grok 3 Mini Beta (high)x-ai/grok-3-mini-beta:high | xAI | 42.6 | — | 5 of 6 benchmarks | 5.9% | 0.0% | 77.8% | 88.1% | 1387 | |
| 272 | QwQ 32Bqwen/qwq-32b | Qwen | 42.6 | — | 1 of 6 benchmarks | 1364 | |||||
| 273 | Gemini 2.5 Flash Lite (nothinking)google/gemini-2.5-flash-lite:nothinking | 42.5 | $0.10 / $0.40 | 1 of 6 benchmarks | 1364 | ||||||
| 274 | Gemini 3.5 Flash Lite (minimal)google/gemini-3.5-flash-lite:minimal | 42.4 | $0.30 / $2.50 | 1 of 6 benchmarks | 51.1% | ||||||
| 275 | GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | 42.4 | $0.40 / $1.60 | 4 of 6 benchmarks | 10.0% | 44.7% | 87.3% | 1354 | ||
| 276 | Qwen 2.5 (max)qwen/qwen-2.5:max | Qwen | 42.3 | — | 1 of 6 benchmarks | 1363 | |||||
| 277 | Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | 42.3 | $1.00 / $5.00 | 4 of 6 benchmarks | 4.1% | 35.8% | 86.9% | 1398 | ||
| 278 | Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | MoonshotAI | 42.3 | $0.60 / $2.50 | 2 of 6 benchmarks | 21.4% | 0.0% | ||||
| 279 | O1 Mini (high)openai/o1-mini:high | OpenAI | 42.2 | — | 3 of 6 benchmarks | 1.4% | 46.9% | 89.2% | |||
| 280 | GLM 4.7z-ai/glm-4.7 | Z.ai | 42.1 | $0.40 / $1.75 | 4 of 6 benchmarks | 2.4% | 0.0% | 83.3% | 1428 | ||
| 281 | Step 3stepfun/step-3 | StepFun | 42.0 | — | 1 of 6 benchmarks | 1362 | |||||
| 282 | O1 Miniopenai/o1-mini | OpenAI | 41.9 | — | 1 of 6 benchmarks | 1362 | |||||
| 283 | Trinity Large Previewarcee-ai/trinity-large-preview | Arcee AI | 41.8 | — | 1 of 6 benchmarks | 1362 | |||||
| 284 | GLM 4.5Vz-ai/glm-4.5v | Z.ai | 41.6 | $0.60 / $1.80 | 1 of 6 benchmarks | 1360 | |||||
| 285 | DeepSeek V4 Pro (none)deepseek/deepseek-v4-pro:none | DeepSeek | 41.4 | $1.168 / $2.336 | 1 of 6 benchmarks | 46.7% | |||||
| 286 | GPT-5 Nano (low)openai/gpt-5-nano:low | OpenAI | 41.4 | $0.05 / $0.40 | 1 of 6 benchmarks | 46.7% | |||||
| 287 | GPT-5 (minimal)openai/gpt-5:minimal | OpenAI | 41.4 | $1.25 / $10.00 | 1 of 6 benchmarks | 46.7% | |||||
| 288 | GPT-5.4 Nano (none)openai/gpt-5.4-nano:none | OpenAI | 41.4 | $0.20 / $1.25 | 1 of 6 benchmarks | 46.7% | |||||
| 289 | MiniMax M2minimax/minimax-m2 | MiniMax | 41.3 | $0.255 / $1.02 | 1 of 6 benchmarks | 1354 | |||||
| 290 | Claude Opus 4 (27K)anthropic/claude-opus-4:27k | Anthropic | 41.3 | $15.00 / $75.00 | 3 of 6 benchmarks | 4.1% | 4.2% | 64.4% | |||
| 291 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Qwen | 41.0 | $0.12 / $0.50 | 1 of 6 benchmarks | 1352 | |||||
| 292 | Ling Flash 2.0inclusionai/ling-flash-2.0 | inclusionAI | 40.9 | — | 1 of 6 benchmarks | 1352 | |||||
| 293 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | NVIDIA | 40.6 | $0.05 / $0.20 | 1 of 6 benchmarks | 1351 | |||||
| 294 | Gemini 3.1 Flash Lite (low)google/gemini-3.1-flash-lite:low | 40.6 | $0.25 / $1.50 | 1 of 6 benchmarks | 44.4% | ||||||
| 295 | o3 Mini (low)openai/o3-mini:low | OpenAI | 40.6 | $1.10 / $4.40 | 1 of 6 benchmarks | 44.4% | |||||
| 296 | Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | 40.3 | $3.00 / $15.00 | 4 of 6 benchmarks | 9.3% | 2.1% | 35.6% | 1427 | ||
| 297 | Hunyuan TurboStencent/hunyuan-turbos | Tencent | 40.3 | — | 1 of 6 benchmarks | 1347 | |||||
| 298 | GPT-5.6 Luna (none)openai/gpt-5.6-luna:none | OpenAI | 40.1 | $0.10 / $0.60 | 1 of 6 benchmarks | 40.0% | |||||
| 299 | Ring Flash 2.0inclusionai/ring-flash-2.0 | inclusionAI | 40.1 | — | 1 of 6 benchmarks | 1340 | |||||
| 300 | Mistral Small 2506mistralai/mistral-small-2506 | Mistral AI | 39.8 | — | 1 of 6 benchmarks | 1338 | |||||
| 301 | O1 Mini (medium)openai/o1-mini:medium | OpenAI | 39.7 | — | 3 of 6 benchmarks | 1.7% | 44.7% | 84.3% | |||
| 302 | Claude 3.7 Sonnet (64K)anthropic/claude-3-7-sonnet:64k | Anthropic | 39.6 | — | 4 of 6 benchmarks | 3.1% | 0.0% | 57.8% | 91.2% | ||
| 303 | gpt-oss-20bopenai/gpt-oss-20b | OpenAI | 39.6 | $0.03 / $0.13 | 1 of 6 benchmarks | 1336 | |||||
| 304 | Nova 2 Liteamazon/nova-2-lite-v1 | Amazon | 39.5 | $0.30 / $2.50 | 1 of 6 benchmarks | 1333 | |||||
| 305 | Gemini 3.1 Flash Lite (minimal)google/gemini-3.1-flash-lite:minimal | 39.5 | $0.25 / $1.50 | 1 of 6 benchmarks | 37.8% | ||||||
| 306 | Gemini 2.0 Flash Lite Preview 02 05google/gemini-2.0-flash-lite-preview-02-05 | 39.3 | — | 1 of 6 benchmarks | 1326 | ||||||
| 307 | Granite 4.1 8Bibm-granite/granite-4.1-8b | IBM | 39.0 | $0.05 / $0.10 | 1 of 6 benchmarks | 1318 | |||||
| 308 | Gemma 3 12Bgoogle/gemma-3-12b-it | 38.9 | $0.05 / $0.15 | 1 of 6 benchmarks | 1318 | ||||||
| 309 | GPT-5 Nano (minimal)openai/gpt-5-nano:minimal | OpenAI | 38.8 | $0.05 / $0.40 | 1 of 6 benchmarks | 35.6% | |||||
| 310 | OLMo 3 32B Thinkallenai/olmo-3-32b-think | Allen Institute for AI | 38.5 | — | 1 of 6 benchmarks | 1311 | |||||
| 311 | Claude Opus 4anthropic/claude-opus-4 | Anthropic | 38.4 | $15.00 / $75.00 | 5 of 6 benchmarks | 4.5% | 0.0% | 42.2% | 85.0% | 1404 | |
| 312 | Command A (03-2025)cohere/command-a-03-2025 | Cohere | 38.2 | — | 1 of 6 benchmarks | 1309 | |||||
| 313 | Grok 3 Betax-ai/grok-3-beta | xAI | 37.9 | — | 5 of 6 benchmarks | 3.8% | 0.0% | 55.6% | 88.8% | 1373 | |
| 314 | GLM 5.2 (none)z-ai/glm-5.2:none | Z.ai | 37.8 | $0.462 / $1.452 | 1 of 6 benchmarks | 28.9% | |||||
| 315 | Claude Sonnet 4 (59K)anthropic/claude-sonnet-4:59k | Anthropic | 37.6 | $3.00 / $15.00 | 2 of 6 benchmarks | 0.0% | 68.9% | ||||
| 316 | OLMo 3.1 32B Instructallenai/olmo-3.1-32b-instruct | Allen Institute for AI | 37.6 | — | 1 of 6 benchmarks | 1304 | |||||
| 317 | GPT-5.4 Mini (none)openai/gpt-5.4-mini:none | OpenAI | 37.5 | $0.75 / $4.50 | 1 of 6 benchmarks | 26.7% | |||||
| 318 | Step 1o Turbo 202506stepfun/step-1o-turbo-202506 | StepFun | 37.2 | — | 1 of 6 benchmarks | 1297 | |||||
| 319 | Qwen3 235B A22Bqwen/qwen3-235b-a22b | Qwen | 36.8 | $0.455 / $1.82 | 3 of 6 benchmarks | 0.0% | 68.9% | 1392 | |||
| 320 | OLMo 3.1 32B Thinkallenai/olmo-3.1-32b-think | Allen Institute for AI | 36.8 | — | 1 of 6 benchmarks | 1296 | |||||
| 321 | Magistral Medium 2506mistralai/magistral-medium-2506 | Mistral AI | 36.0 | — | 1 of 6 benchmarks | 1285 | |||||
| 322 | GPT-4.1openai/gpt-4.1 | OpenAI | 35.9 | $2.00 / $8.00 | 5 of 6 benchmarks | 5.5% | 0.0% | 38.3% | 83.0% | 1374 | |
| 323 | Gemma 3 27Bgoogle/gemma-3-27b-it | 35.8 | $0.08 / $0.45 | 3 of 6 benchmarks | 22.2% | 74.0% | 1323 | ||||
| 324 | Hunyuan Large Visiontencent/hunyuan-large-vision | Tencent | 35.8 | — | 1 of 6 benchmarks | 1280 | |||||
| 325 | Gemini 1.5 Pro 002google/gemini-1.5-pro-002 | 35.7 | — | 3 of 6 benchmarks | 23.1% | 70.4% | 1339 | ||||
| 326 | Claude Sonnet 4anthropic/claude-sonnet-4 | Anthropic | 35.7 | $3.00 / $15.00 | 5 of 6 benchmarks | 4.1% | 0.0% | 28.9% | 84.4% | 1389 | |
| 327 | Gemini 2.0 Flash 001google/gemini-2.0-flash-001 | 35.7 | — | 4 of 6 benchmarks | 1.7% | 31.1% | 82.2% | 1356 | |||
| 328 | Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct | Qwen | 35.2 | — | 1 of 6 benchmarks | 1270 | |||||
| 329 | Claude Opus 4.1 (16K)anthropic/claude-opus-4.1:16k | Anthropic | 34.9 | $15.00 / $75.00 | 2 of 6 benchmarks | 0.0% | 64.4% | ||||
| 330 | Amazon Nova Pro v1.0amazon/amazon-nova-pro-v1.0 | Amazon | 34.9 | — | 1 of 6 benchmarks | 1269 | |||||
| 331 | Gemma 3n E4Bgoogle/gemma-3n-e4b-it | 34.6 | — | 1 of 6 benchmarks | 1260 | ||||||
| 332 | GPT-5.1 (none)openai/gpt-5.1:none | OpenAI | 34.5 | $1.25 / $10.00 | 2 of 6 benchmarks | 2.1% | 37.8% | ||||
| 333 | Gemma 3 4Bgoogle/gemma-3-4b-it | 34.3 | — | 1 of 6 benchmarks | 1254 | ||||||
| 334 | Amazon Nova Lite v1.0amazon/amazon-nova-lite-v1.0 | Amazon | 34.0 | — | 1 of 6 benchmarks | 1244 | |||||
| 335 | Claude 3.7 Sonnetanthropic/claude-3-7-sonnet | Anthropic | 33.9 | — | 4 of 6 benchmarks | 3.1% | 21.9% | 68.2% | 1363 | ||
| 336 | Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | 33.8 | $2.50 / $10.00 | 1 of 6 benchmarks | 1231 | |||||
| 337 | DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | DeepSeek | 33.6 | $0.27 / $1.12 | 4 of 6 benchmarks | 0.0% | 37.8% | 75.5% | 1369 | ||
| 338 | Mistral Medium 2505mistralai/mistral-medium-2505 | Mistral AI | 33.6 | — | 4 of 6 benchmarks | 0.3% | 32.2% | 81.6% | 1348 | ||
| 339 | OLMo 2 0325 32B Instructallenai/olmo-2-0325-32b-instruct | Allen Institute for AI | 33.3 | — | 1 of 6 benchmarks | 1227 | |||||
| 340 | DeepSeek V3deepseek/deepseek-chat | DeepSeek | 33.2 | $0.2574 / $1.0287 | 4 of 6 benchmarks | 1.7% | 48.9% | 64.8% | 1311 | ||
| 341 | Claude Opus 4.1 (32K)anthropic/claude-opus-4.1:32k | Anthropic | 33.2 | $15.00 / $75.00 | 2 of 6 benchmarks | 12.6% | 2.4% | ||||
| 342 | Amazon Nova Micro v1.0amazon/amazon-nova-micro-v1.0 | Amazon | 33.2 | — | 1 of 6 benchmarks | 1224 | |||||
| 343 | QwQ 32B Previewqwen/qwq-32b-preview | Qwen | 33.0 | — | 1 of 6 benchmarks | 1210 | |||||
| 344 | Command R (08-2024)cohere/command-r-08-2024 | Cohere | 32.9 | $0.15 / $0.60 | 1 of 6 benchmarks | 1207 | |||||
| 345 | Gemini 3.5 Flash Lite (high)google/gemini-3.5-flash-lite:high | 32.9 | $0.30 / $2.50 | 3 of 6 benchmarks | 26.0% | 0.0% | 71.1% | ||||
| 346 | Mistral Nemomistralai/mistral-nemo | Mistral | 32.8 | $0.019 / $0.03 | 1 of 6 benchmarks | 10.8% | |||||
| 347 | GLM 4.5z-ai/glm-4.5 | Z.ai | 32.7 | $0.60 / $2.20 | 3 of 6 benchmarks | 0.0% | 0.0% | 1413 | |||
| 348 | Qwen (max)qwen/qwen:max | Qwen | 32.6 | — | 3 of 6 benchmarks | 1.0% | 16.1% | 67.2% | |||
| 349 | Grok 2 1212x-ai/grok-2-1212 | xAI | 31.5 | — | 3 of 6 benchmarks | 0.7% | 11.5% | 63.5% | |||
| 350 | GPT 4 1106 Previewopenai/gpt-4-1106-preview | OpenAI | 31.3 | — | 2 of 6 benchmarks | 40.0% | 1303 | ||||
| 351 | Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct | Meta | 31.1 | — | 4 of 6 benchmarks | 0.7% | 20.6% | 73.0% | 1317 | ||
| 352 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Qwen | 31.1 | $0.36 / $0.40 | 3 of 6 benchmarks | 8.1% | 63.2% | 1296 | |||
| 353 | GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | 29.8 | $0.10 / $0.40 | 4 of 6 benchmarks | 1.0% | 28.9% | 70.0% | 1274 | ||
| 354 | Magistral Small 2506mistralai/magistral-small-2506 | Mistral AI | 29.3 | — | 2 of 6 benchmarks | 0.0% | 30.0% | ||||
| 355 | GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | 29.2 | $5.00 / $15.00 | 3 of 6 benchmarks | 6.3% | 51.0% | 1305 | |||
| 356 | GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | 28.5 | $0.15 / $0.60 | 3 of 6 benchmarks | 6.9% | 52.6% | 1276 | |||
| 357 | GPT-4 Turboopenai/gpt-4-turbo | OpenAI | 27.9 | $10.00 / $30.00 | 3 of 6 benchmarks | 6.7% | 46.7% | 1296 | |||
| 358 | Mistral Large 2407mistralai/mistral-large-2407 | Mistral | 27.4 | $2.00 / $6.00 | 3 of 6 benchmarks | 8.5% | 44.8% | 1288 | |||
| 359 | GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | 27.2 | $2.50 / $10.00 | 3 of 6 benchmarks | 0.3% | 6.3% | 49.8% | |||
| 360 | Mistral Small 3.1 24B Instruct 2503mistralai/mistral-small-3.1-24b-instruct-2503 | Mistral AI | 26.9 | — | 3 of 6 benchmarks | 5.8% | 46.8% | 1278 | |||
| 361 | Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | 26.9 | $2.00 / $6.00 | 2 of 6 benchmarks | 24.2% | 1228 | ||||
| 362 | Gemini 1.5 Pro 001google/gemini-1.5-pro-001 | 26.7 | — | 3 of 6 benchmarks | 6.8% | 40.8% | 1299 | ||||
| 363 | Claude 3.5 Sonnetanthropic/claude-3-5-sonnet | Anthropic | 26.7 | — | 5 of 6 benchmarks | 1.0% | 0.0% | 8.5% | 57.0% | 1351 | |
| 364 | Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct | Meta | 26.5 | — | 4 of 6 benchmarks | 0.0% | 7.8% | 62.3% | 1308 | ||
| 365 | GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | OpenAI | 26.4 | $2.50 / $10.00 | 4 of 6 benchmarks | 0.3% | 6.4% | 53.3% | 1309 | ||
| 366 | Claude 3 Opusanthropic/claude-3-opus | Anthropic | 26.2 | — | 3 of 6 benchmarks | 4.7% | 37.5% | 1312 | |||
| 367 | Gemini 1.5 Flash 002google/gemini-1.5-flash-002 | 26.1 | — | 4 of 6 benchmarks | 0.0% | 16.3% | 61.9% | 1289 | |||
| 368 | Gemini 1.5 Flash 8B 001google/gemini-1.5-flash-8b-001 | 26.0 | — | 2 of 6 benchmarks | 4.6% | 1230 | |||||
| 369 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Meta | 25.9 | $0.10 / $0.32 | 3 of 6 benchmarks | 5.1% | 41.6% | 1296 | |||
| 370 | Mistral Small 24B Instruct 2501mistralai/mistral-small-24b-instruct-2501 | Mistral AI | 25.3 | — | 3 of 6 benchmarks | 5.3% | 44.8% | 1262 | |||
| 371 | Claude 2anthropic/claude-2 | Anthropic | 25.1 | — | 2 of 6 benchmarks | 2.5% | 11.7% | ||||
| 372 | Mistral Large 2411mistralai/mistral-large-2411 | Mistral AI | 25.0 | — | 4 of 6 benchmarks | 0.3% | 7.8% | 50.3% | 1282 | ||
| 373 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Meta | 23.4 | $0.40 / $0.40 | 3 of 6 benchmarks | 3.6% | 36.7% | 1269 | |||
| 374 | Claude 3.5 Haikuanthropic/claude-3-5-haiku | Anthropic | 23.2 | — | 4 of 6 benchmarks | 0.3% | 4.3% | 46.4% | 1286 | ||
| 375 | Gemini 1.5 Flash 001google/gemini-1.5-flash-001 | 22.8 | — | 3 of 6 benchmarks | 3.9% | 25.1% | 1258 | ||||
| 376 | Claude 3 Sonnetanthropic/claude-3-sonnet | Anthropic | 21.5 | — | 3 of 6 benchmarks | 2.5% | 18.2% | 1253 | |||
| 377 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Meta | 21.0 | $0.05 / $0.08 | 3 of 6 benchmarks | 2.5% | 22.9% | 1189 | |||
| 378 | Claude 3 Haikuanthropic/claude-3-haiku | Anthropic | 20.9 | $0.25 / $1.25 | 3 of 6 benchmarks | 1.8% | 14.9% | 1231 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- Epoch AI Benchmarking HubCC BY 4.0
Benchmark runs by Epoch AI, from the AI Benchmarking Hub.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.