---
title: "Best LLM for vision"
description: "Models ranked across the LMArena vision arenas: OCR, diagrams, homework, entity recognition and captioning, with coverage shown for every row."
url: "https://aldena.ai/best-llm-for-vision"
---

# Best LLM for vision

Vision covers reading a document, following a diagram, and describing what is in a frame.

Aldena runs these models inside your team rooms. [See what each one costs](https://aldena.ai/models).

| rank | model | vendor | score | coverage (benchmarks) | price (in / out) | LMArena vision (elo) | LMArena vision OCR (elo) | LMArena vision diagrams (elo) | LMArena vision homework (elo) | LMArena vision entity recognition (elo) | LMArena vision captioning (elo) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Fable 5 | Anthropic | 84.0 | 4 of 6 | $10.00 / $50.00 | 1315 | 1331 | 1363 | 1346 | — | — |
| 2 | Qwen3.8 Max | Qwen | 81.7 | 4 of 6 | $2.00 / $6.00 | 1301 | 1317 | 1327 | 1331 | — | — |
| 3 | Claude Opus 4.6 (thinking) | Anthropic | 81.3 | 5 of 6 | $5.00 / $25.00 | 1300 | 1315 | 1325 | 1325 | 1269 | — |
| 4 | Claude Opus 5 (high) | Anthropic | 81.0 | 4 of 6 | $5.00 / $25.00 | 1297 | 1308 | 1328 | 1344 | — | — |
| 5 | Claude Opus 4.7 | Anthropic | 80.9 | 4 of 6 | $5.00 / $25.00 | 1299 | 1314 | 1333 | 1327 | — | — |
| 6 | Claude Opus 4.7 (thinking) | Anthropic | 80.5 | 5 of 6 | $5.00 / $25.00 | 1301 | 1313 | 1337 | 1330 | 1234 | — |
| 7 | Gemini 3 Pro | Google | 79.9 | 6 of 6 | — | 1289 | 1303 | 1300 | 1309 | 1299 | 1271 |
| 8 | Grok 4.5 | SpaceXAI | 78.3 | 4 of 6 | $2.00 / $6.00 | 1285 | 1296 | 1328 | 1337 | — | — |
| 9 | Gemini 3.1 Pro Preview | Google | 77.9 | 6 of 6 | $2.00 / $12.00 | 1277 | 1294 | 1309 | 1311 | 1286 | 1274 |
| 10 | Claude Opus 4.8 (thinking) | Anthropic | 77.9 | 4 of 6 | $5.00 / $25.00 | 1284 | 1299 | 1320 | 1338 | — | — |
| 11 | GPT-5.5 | OpenAI | 77.6 | 4 of 6 | $5.00 / $30.00 | 1286 | 1299 | 1326 | 1330 | — | — |
| 12 | GPT-5.5 (high) | OpenAI | 76.9 | 4 of 6 | $5.00 / $30.00 | 1283 | 1299 | 1314 | 1339 | — | — |
| 13 | Claude Opus 4.6 | Anthropic | 75.9 | 5 of 6 | $5.00 / $25.00 | 1293 | 1310 | 1331 | 1314 | 1222 | — |
| 14 | Gemini 3.6 Flash | Google | 75.1 | 3 of 6 | $0.75 / $3.75 | 1295 | 1308 | 1313 | — | — | — |
| 15 | Muse Spark | Meta | 74.4 | 4 of 6 | — | 1294 | 1302 | 1313 | 1300 | — | — |
| 16 | Muse Spark 1.2 (xhigh) | Meta | 74.4 | 3 of 6 | $1.25 / $4.25 | 1290 | 1299 | 1316 | — | — | — |
| 17 | Gemini 3.5 Flash (medium) | Google | 72.8 | 4 of 6 | $1.50 / $9.00 | 1284 | 1296 | 1297 | 1322 | — | — |
| 18 | GPT-5.4 | OpenAI | 72.6 | 4 of 6 | $2.50 / $15.00 | 1280 | 1294 | 1318 | 1307 | — | — |
| 19 | Gemini 3 Flash Preview | Google | 72.4 | 6 of 6 | $0.50 / $3.00 | 1271 | 1285 | 1294 | 1307 | 1293 | 1226 |
| 20 | Claude Sonnet 5 (high) | Anthropic | 72.0 | 4 of 6 | $2.00 / $10.00 | 1273 | 1291 | 1320 | 1313 | — | — |
| 21 | Claude Opus 4.8 | Anthropic | 71.6 | 4 of 6 | $5.00 / $25.00 | 1280 | 1293 | 1304 | 1322 | — | — |
| 22 | GPT-5.6 Sol (xhigh) | OpenAI | 70.7 | 4 of 6 | $5.00 / $30.00 | 1280 | 1288 | 1308 | 1313 | — | — |
| 23 | GPT-5.4 (high) | OpenAI | 70.1 | 5 of 6 | $2.50 / $15.00 | 1283 | 1299 | 1319 | 1330 | 1183 | — |
| 24 | GPT-5.6 Terra (xhigh) | OpenAI | 69.8 | 4 of 6 | $1.00 / $6.00 | 1270 | 1280 | 1312 | 1317 | — | — |
| 25 | Gemini 3.5 Flash (high) | Google | 69.6 | 4 of 6 | $1.50 / $9.00 | 1283 | 1292 | 1304 | 1294 | — | — |
| 26 | Gemini 3 Flash Preview (thinking minimal) | Google | 68.7 | 6 of 6 | $0.50 / $3.00 | 1259 | 1271 | 1284 | 1298 | 1278 | 1213 |
| 27 | GPT 5.5 Instant | OpenAI | 67.7 | 4 of 6 | — | 1278 | 1286 | 1315 | 1282 | — | — |
| 28 | GPT-5.2 Chat | OpenAI | 67.3 | 5 of 6 | $1.75 / $14.00 | 1278 | 1288 | 1310 | 1293 | 1223 | — |
| 29 | Muse Spark 1.1 | Meta | 67.3 | 4 of 6 | $1.25 / $4.25 | 1282 | 1293 | 1299 | 1279 | — | — |
| 30 | Kimi K2.6 | MoonshotAI | 66.1 | 4 of 6 | $0.5415 / $2.28 | 1263 | 1278 | 1292 | 1302 | — | — |
| 31 | GPT-5.1 (high) | OpenAI | 64.7 | 6 of 6 | $1.25 / $10.00 | 1250 | 1258 | 1275 | 1282 | 1240 | 1240 |
| 32 | Gemini 2.5 Pro | Google | 63.8 | 6 of 6 | $1.25 / $10.00 | 1246 | 1258 | 1266 | 1272 | 1252 | 1251 |
| 33 | Gemini 3.5 Flash Lite | Google | 63.2 | 3 of 6 | $0.30 / $2.50 | 1260 | 1282 | 1283 | — | — | — |
| 34 | Dola Seed 2.0 Pro | ByteDance | 62.0 | 4 of 6 | — | 1258 | 1270 | 1287 | 1279 | — | — |
| 35 | Qwen3.7 Plus | Qwen | 61.6 | 4 of 6 | $0.32 / $1.28 | 1262 | 1278 | 1279 | 1276 | — | — |
| 36 | Grok 4.20 Beta 0309 (reasoning) | xAI | 61.4 | 5 of 6 | — | 1255 | 1264 | 1273 | 1263 | 1247 | — |
| 37 | Claude Sonnet 4.6 | Anthropic | 61.3 | 5 of 6 | $3.00 / $15.00 | 1275 | 1291 | 1306 | 1301 | 1159 | — |
| 38 | Kimi K2.5 (thinking) | MoonshotAI | 60.8 | 6 of 6 | $0.57 / $2.85 | 1249 | 1264 | 1277 | 1290 | 1230 | 1196 |
| 39 | Qwen3.5 397B A17B | Qwen | 60.2 | 5 of 6 | $0.39 / $2.34 | 1247 | 1262 | 1276 | 1283 | 1226 | — |
| 40 | Gemini 3.1 Flash Lite Preview | Google | 60.0 | 6 of 6 | $0.25 / $1.50 | 1234 | 1247 | 1251 | 1270 | 1272 | 1226 |
| 41 | GPT-5.6 Luna (xhigh) | OpenAI | 59.9 | 4 of 6 | $0.10 / $0.60 | 1249 | 1256 | 1281 | 1285 | — | — |
| 42 | GPT-5.2 (high) | OpenAI | 59.7 | 6 of 6 | $1.75 / $14.00 | 1244 | 1260 | 1274 | 1284 | 1185 | 1268 |
| 43 | Grok 4.20 Multi Agent Beta 0309 | xAI | 59.6 | 5 of 6 | — | 1252 | 1261 | 1268 | 1273 | 1231 | — |
| 44 | GPT-5.4 Mini (high) | OpenAI | 57.1 | 5 of 6 | $0.75 / $4.50 | 1253 | 1266 | 1284 | 1293 | 1182 | — |
| 45 | Grok 4.3 | SpaceXAI | 55.6 | 5 of 6 | $1.25 / $2.50 | 1244 | 1253 | 1268 | 1256 | 1227 | — |
| 46 | Gemma 4 31B | Google | 55.3 | 6 of 6 | $0.10 / $0.34 | 1256 | 1270 | 1281 | 1300 | 1191 | 1165 |
| 47 | Gemma 4 26B A4B  | Google | 54.6 | 6 of 6 | $0.12 / $0.40 | 1240 | 1256 | 1261 | 1285 | 1153 | 1248 |
| 48 | Chatgpt 4o | OpenAI | 54.5 | 6 of 6 | — | 1241 | 1249 | 1264 | 1247 | 1225 | 1212 |
| 49 | MiniMax M3 | MiniMax | 53.7 | 4 of 6 | $0.30 / $1.20 | 1240 | 1253 | 1263 | 1263 | — | — |
| 50 | GPT 4.5 Preview (partial coverage) | OpenAI | 52.5 | 1 of 6 | — | 1226 | — | — | — | — | — |
| 51 | GPT-5 (high) | OpenAI | 52.4 | 6 of 6 | $1.25 / $10.00 | 1213 | 1232 | 1252 | 1260 | 1257 | 1191 |
| 52 | GPT 5 Chat | OpenAI | 50.6 | 6 of 6 | — | 1225 | 1245 | 1260 | 1274 | 1188 | 1198 |
| 53 | Kimi K2.5 Instant | Moonshot AI | 49.0 | 4 of 6 | — | 1237 | 1245 | 1249 | 1247 | — | — |
| 54 | MiMo-V2.5 | Xiaomi | 48.8 | 5 of 6 | $0.14 / $0.28 | 1238 | 1253 | 1271 | 1260 | 1172 | — |
| 55 | Gemini 2.5 Flash | Google | 48.4 | 6 of 6 | $0.30 / $2.50 | 1214 | 1222 | 1233 | 1238 | 1220 | 1215 |
| 56 | Qwen3.5-122B-A10B | Qwen | 48.1 | 4 of 6 | $0.29 / $2.40 | 1227 | 1239 | 1246 | 1252 | — | — |
| 57 | GPT-5.1 | OpenAI | 47.9 | 6 of 6 | $1.25 / $10.00 | 1238 | 1252 | 1264 | 1254 | 1188 | 1180 |
| 58 | Qwen3 VL 235B A22B Instruct | Qwen | 47.7 | 6 of 6 | $0.26 / $1.04 | 1215 | 1229 | 1239 | 1260 | 1190 | 1208 |
| 59 | o1 (partial coverage) | OpenAI | 47.2 | 1 of 6 | $15.00 / $60.00 | 1193 | — | — | — | — | — |
| 60 | o3 | OpenAI | 46.8 | 6 of 6 | $2.00 / $8.00 | 1217 | 1223 | 1231 | 1247 | 1234 | 1181 |
| 61 | GPT-5.2 | OpenAI | 46.5 | 6 of 6 | $1.75 / $14.00 | 1229 | 1238 | 1257 | 1267 | 1192 | 1167 |
| 62 | Mimo v2 Omni | Xiaomi | 46.5 | 4 of 6 | — | 1217 | 1231 | 1254 | 1246 | — | — |
| 63 | GLM 5V Turbo | Z.ai | 46.4 | 5 of 6 | $1.20 / $4.00 | 1231 | 1244 | 1264 | 1254 | 1181 | — |
| 64 | Gemini 1.5 Pro 002 (partial coverage) | Google | 44.9 | 1 of 6 | — | 1180 | — | — | — | — | — |
| 65 | GPT-4.1 | OpenAI | 44.2 | 6 of 6 | $2.00 / $8.00 | 1214 | 1226 | 1232 | 1249 | 1183 | 1207 |
| 66 | GPT-4o (2024-05-13) (partial coverage) | OpenAI | 43.5 | 1 of 6 | $5.00 / $15.00 | 1162 | — | — | — | — | — |
| 67 | Ernie 5.0 Preview 1220 | Baidu | 43.4 | 4 of 6 | — | 1218 | 1230 | 1223 | 1237 | — | — |
| 68 | Claude Sonnet 4 (thinking 32K) | Anthropic | 43.0 | 3 of 6 | $3.00 / $15.00 | 1208 | 1216 | 1228 | — | — | — |
| 69 | o4 Mini | OpenAI | 42.5 | 6 of 6 | $1.10 / $4.40 | 1202 | 1210 | 1220 | 1251 | 1218 | 1182 |
| 70 | Qwen VL (max) | Qwen | 42.0 | 3 of 6 | — | 1186 | 1214 | 1249 | — | — | — |
| 71 | Qwen3.5-27B | Qwen | 40.8 | 5 of 6 | $0.195 / $1.56 | 1219 | 1232 | 1242 | 1242 | 1149 | — |
| 72 | Mistral Large 3 | Mistral AI | 40.4 | 4 of 6 | — | 1199 | 1219 | 1241 | 1221 | — | — |
| 73 | GPT-4.1 Mini | OpenAI | 40.2 | 6 of 6 | $0.40 / $1.60 | 1203 | 1206 | 1209 | 1222 | 1198 | 1195 |
| 74 | Gemini 1.5 Flash 002 (partial coverage) | Google | 40.2 | 1 of 6 | — | 1141 | — | — | — | — | — |
| 75 | Mistral Medium 3.5 | Mistral | 39.8 | 4 of 6 | $1.50 / $7.50 | 1198 | 1215 | 1242 | 1216 | — | — |
| 76 | Gemini 2.0 Flash Lite Preview 02 05 (partial coverage) | Google | 39.6 | 1 of 6 | — | 1136 | — | — | — | — | — |
| 77 | Qwen2.5 VL 72B Instruct (partial coverage) | Qwen | 38.5 | 1 of 6 | — | 1122 | — | — | — | — | — |
| 78 | Gemini 1.5 Pro 001 (partial coverage) | Google | 38.2 | 1 of 6 | — | 1120 | — | — | — | — | — |
| 79 | Claude Opus 4 (thinking 16K) | Anthropic | 38.1 | 4 of 6 | $15.00 / $75.00 | 1207 | 1216 | 1203 | 1216 | — | — |
| 80 | Qwen2.5 VL 32B Instruct (partial coverage) | Qwen | 37.9 | 1 of 6 | — | 1119 | — | — | — | — | — |
| 81 | GPT-4o (2024-08-06) (partial coverage) | OpenAI | 37.6 | 1 of 6 | $2.50 / $10.00 | 1119 | — | — | — | — | — |
| 82 | GPT-5 Mini (high) | OpenAI | 37.6 | 6 of 6 | $0.25 / $2.00 | 1183 | 1198 | 1214 | 1233 | 1207 | 1168 |
| 83 | GPT-4 Turbo (partial coverage) | OpenAI | 37.4 | 1 of 6 | $10.00 / $30.00 | 1112 | — | — | — | — | — |
| 84 | Claude 3.7 Sonnet (thinking 32K) | Anthropic | 37.1 | 4 of 6 | — | 1196 | 1210 | 1216 | 1210 | — | — |
| 85 | GPT-4o-mini (2024-07-18) (partial coverage) | OpenAI | 36.8 | 1 of 6 | $0.15 / $0.60 | 1098 | — | — | — | — | — |
| 86 | Qwen3 VL 235B A22B Thinking | Qwen | 36.7 | 4 of 6 | $0.40 / $4.00 | 1190 | 1201 | 1210 | 1227 | — | — |
| 87 | Claude Opus 4 | Anthropic | 36.7 | 4 of 6 | $15.00 / $75.00 | 1189 | 1197 | 1203 | 1237 | — | — |
| 88 | GPT-4.1 Nano (partial coverage) | OpenAI | 36.5 | 1 of 6 | $0.10 / $0.40 | 1089 | — | — | — | — | — |
| 89 | Grok 4.1 Fast (reasoning) | xAI | 36.4 | 5 of 6 | — | 1195 | 1192 | 1215 | 1170 | 1202 | — |
| 90 | Gemini 1.5 Flash 8B 001 (partial coverage) | Google | 36.3 | 1 of 6 | — | 1072 | — | — | — | — | — |
| 91 | Claude 3 Opus (partial coverage) | Anthropic | 36.0 | 1 of 6 | — | 1062 | — | — | — | — | — |
| 92 | Gemini 1.5 Flash 001 (partial coverage) | Google | 35.7 | 1 of 6 | — | 1060 | — | — | — | — | — |
| 93 | Grok 4 0709 | xAI | 35.6 | 6 of 6 | — | 1182 | 1175 | 1191 | 1169 | 1236 | 1167 |
| 94 | Amazon Nova Pro v1.0 (partial coverage) | Amazon | 35.4 | 1 of 6 | — | 1019 | — | — | — | — | — |
| 95 | Amazon Nova Lite v1.0 (partial coverage) | Amazon | 35.1 | 1 of 6 | — | 1019 | — | — | — | — | — |
| 96 | Claude Sonnet 4 | Anthropic | 35.1 | 4 of 6 | $3.00 / $15.00 | 1189 | 1190 | 1202 | 1222 | — | — |
| 97 | Claude 3 Sonnet (partial coverage) | Anthropic | 34.8 | 1 of 6 | — | 1017 | — | — | — | — | — |
| 98 | GPT-5.4 Nano (high) | OpenAI | 34.7 | 5 of 6 | $0.20 / $1.25 | 1202 | 1215 | 1229 | 1235 | 1119 | — |
| 99 | Claude 3 Haiku (partial coverage) | Anthropic | 34.6 | 1 of 6 | $0.25 / $1.25 | 1001 | — | — | — | — | — |
| 100 | Gemini 2.5 Flash Lite (thinking) | Google | 33.9 | 6 of 6 | $0.10 / $0.40 | 1188 | 1188 | 1187 | 1190 | 1185 | 1185 |
| 101 | Claude 3.7 Sonnet | Anthropic | 33.9 | 4 of 6 | — | 1176 | 1186 | 1192 | 1228 | — | — |
| 102 | Hunyuan Vision 1.5 (thinking) | Tencent | 31.3 | 4 of 6 | — | 1159 | 1161 | 1186 | 1229 | — | — |
| 103 | Gemini 2.5 Flash Lite (nothinking) | Google | 31.2 | 4 of 6 | $0.10 / $0.40 | 1174 | 1185 | 1183 | 1194 | — | — |
| 104 | Claude 3.5 Sonnet | Anthropic | 29.2 | 4 of 6 | — | 1161 | 1177 | 1184 | 1170 | — | — |
| 105 | GLM 4.6V | Z.ai | 27.8 | 4 of 6 | $0.30 / $0.90 | 1164 | 1171 | 1186 | 1155 | — | — |
| 106 | Gemini 2.0 Flash 001 | Google | 27.7 | 5 of 6 | — | 1172 | 1166 | 1176 | 1190 | — | 1148 |
| 107 | GPT-5 Nano (high) | OpenAI | 26.1 | 4 of 6 | $0.05 / $0.40 | 1148 | 1158 | 1153 | 1187 | — | — |
| 108 | Hunyuan Large Vision | Tencent | 26.0 | 3 of 6 | — | 1150 | 1146 | 1114 | — | — | — |
| 109 | GLM 4.5V | Z.ai | 25.9 | 4 of 6 | $0.60 / $1.80 | 1156 | 1158 | 1169 | 1170 | — | — |
| 110 | Step 3 | StepFun | 24.4 | 4 of 6 | — | 1145 | 1151 | 1146 | 1176 | — | — |
| 111 | Gemma 3 27B | Google | 24.0 | 6 of 6 | $0.08 / $0.45 | 1160 | 1162 | 1166 | 1172 | 1175 | 1118 |
| 112 | Step 1o Turbo 202506 | StepFun | 23.9 | 4 of 6 | — | 1157 | 1154 | 1162 | 1138 | — | — |
| 113 | Llama 4 Maverick 17B 128e Instruct | Meta | 23.6 | 4 of 6 | — | 1147 | 1153 | 1152 | 1163 | — | — |
| 114 | Mistral Medium 2508 | Mistral AI | 23.0 | 6 of 6 | — | 1159 | 1172 | 1178 | 1170 | 1132 | 1130 |
| 115 | Molmo 2 8B | Allen Institute for AI | 22.3 | 3 of 6 | — | 1108 | 1087 | 1108 | — | — | — |
| 116 | Llama 4 Scout 17B 16e Instruct | Meta | 21.1 | 4 of 6 | — | 1128 | 1130 | 1148 | 1149 | — | — |
| 117 | Claude 3.5 Haiku | Anthropic | 20.7 | 4 of 6 | — | 1128 | 1129 | 1135 | 1150 | — | — |
| 118 | Mistral Medium 2505 | Mistral AI | 20.0 | 6 of 6 | — | 1156 | 1159 | 1167 | 1165 | 1123 | 1092 |
| 119 | Mistral Small 2506 | Mistral AI | 19.0 | 6 of 6 | — | 1141 | 1142 | 1140 | 1176 | 1134 | 1055 |
| 120 | Mistral Small 3.1 24B Instruct 2503 | Mistral AI | 17.9 | 6 of 6 | — | 1128 | 1133 | 1130 | 1159 | 1127 | 1135 |

## How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

## Data sources

- [LMArena](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset): CC BY 4.0. Arena ratings by LMArena, from the public leaderboard dataset.
