---
title: "Best AI for transcription"
description: "Speech recognition ranked on Open ASR Leaderboard word error rates across English, long-form and five European languages, plus RTFx speed."
url: "https://aldena.ai/best-ai-for-transcription"
---

# Best AI for transcription

Word error rate is lower-is-better and inverted before ranking. RTFx is carried as a separate speed benchmark, never folded into accuracy.

Aldena runs these models inside your team rooms. [See what each one costs](https://aldena.ai/models).

| rank | model | vendor | score | coverage (benchmarks) | Open ASR English word error rate (% wer, lower is better) | Open ASR long-form word error rate (% wer, lower is better) | Open ASR French word error rate (% wer, lower is better) | Open ASR German word error rate (% wer, lower is better) | Open ASR Spanish word error rate (% wer, lower is better) | Open ASR Italian word error rate (% wer, lower is better) | Open ASR Portuguese word error rate (% wer, lower is better) | Open ASR speed (RTFx) (rtfx) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Scribe v2 | ElevenLabs | 86.4 | 7 of 8 | 4.7% | 7.3% | 3.3% | 2.3% | 2.3% | 2.5% | 2.9% | — |
| 2 | Azure Speech | Microsoft | 85.8 | 6 of 8 | 4.5% | — | 3.3% | 1.9% | 2.3% | 2.5% | 3.1% | — |
| 3 | Resonant 1 | Reson8 | 79.7 | 7 of 8 | 4.6% | 8.5% | 4.1% | 2.8% | 2.7% | 3.0% | 3.5% | — |
| 4 | Resonant 1 Flash | Reson8 | 79.4 | 7 of 8 | 4.7% | 8.4% | 4.1% | 2.8% | 2.7% | 3.0% | 3.5% | — |
| 5 | Universal 3 Pro | AssemblyAI | 78.0 | 7 of 8 | 5.2% | 8.3% | 3.8% | 2.6% | 2.5% | 4.1% | 3.9% | — |
| 6 | Cohere Transcribe | Cohere | 68.4 | 8 of 8 | 5.2% | 9.7% | 4.0% | 3.1% | 2.8% | 3.0% | 6.0% | 915.6x |
| 7 | Multilingual | Modulate | 67.4 | 5 of 8 | — | — | 4.1% | 3.5% | 3.0% | 3.2% | 4.9% | — |
| 8 | Vfast (partial coverage) | Modulate | 65.5 | 1 of 8 | 4.4% | — | — | — | — | — | — | — |
| 9 | Parakeet Tdt 0.6B v2 | NVIDIA | 63.5 | 3 of 8 | 5.4% | 11.2% | — | — | — | — | — | 6038.1x |
| 10 | Enhanced | Speechmatics | 62.5 | 7 of 8 | 5.9% | 8.8% | 4.9% | 2.8% | 2.7% | 5.4% | 5.0% | — |
| 11 | Scribe v1 (partial coverage) | Zoom | 62.5 | 1 of 8 | 4.7% | — | — | — | — | — | — | — |
| 12 | Voxtral Small 24B 2507 | Mistral | 61.1 | 7 of 8 | 5.7% | — | 4.1% | 2.9% | 2.9% | 3.8% | 4.4% | 100.1x |
| 13 | STT Async v5 | Soniox | 60.4 | 5 of 8 | — | — | 6.2% | 3.6% | 3.4% | 4.2% | 4.0% | — |
| 14 | Fusion (partial coverage) | Rev | 58.8 | 1 of 8 | — | 9.5% | — | — | — | — | — | — |
| 15 | Parakeet TDT 0.6B v3 | NVIDIA | 58.7 | 8 of 8 | 5.7% | 10.7% | 5.4% | 4.1% | 3.7% | 4.7% | 6.1% | 6098.2x |
| 16 | Universal 3.5 Pro (partial coverage) | AssemblyAI | 58.6 | 1 of 8 | 5.0% | — | — | — | — | — | — | — |
| 17 | Canary Qwen 2.5B | NVIDIA | 57.9 | 3 of 8 | 5.1% | 11.2% | — | — | — | — | — | 861.5x |
| 18 | Machine (partial coverage) | Rev | 57.5 | 1 of 8 | — | 9.6% | — | — | — | — | — | — |
| 19 | ARK ASR 3B | AutoArk AI | 56.4 | 2 of 8 | 4.6% | — | — | — | — | — | — | 484.3x |
| 20 | Parakeet Tdt CTC (110M) | NVIDIA | 55.6 | 2 of 8 | 6.6% | — | — | — | — | — | — | 6118.8x |
| 21 | Granite Speech 4.1 2B | IBM | 55.6 | 2 of 8 | 4.9% | — | — | — | — | — | — | 546.8x |
| 22 | Granite Speech 4.1 2B NAR | IBM | 55.4 | 6 of 8 | 5.0% | — | 5.4% | 4.3% | 3.6% | — | 6.5% | 2079.3x |
| 23 | Moonshine Streaming (medium) | Useful Sensors | 55.3 | 2 of 8 | 5.8% | — | — | — | — | — | — | 2681.0x |
| 24 | Avalon v1 EN (partial coverage) | Aqua Voice | 54.9 | 1 of 8 | 5.2% | — | — | — | — | — | — | — |
| 25 | Canary 1B Flash | NVIDIA | 54.1 | 2 of 8 | 5.8% | — | — | — | — | — | — | 2126.1x |
| 26 | Canary 1B v2 | NVIDIA | 54.0 | 7 of 8 | 6.4% | — | 4.8% | 4.1% | 3.2% | 4.8% | 6.3% | 1821.4x |
| 27 | Solaria 3 (partial coverage) | Upstage | 54.0 | 1 of 8 | 5.3% | — | — | — | — | — | — | — |
| 28 | ARK ASR 0.6B | AutoArk AI | 54.0 | 2 of 8 | 5.1% | — | — | — | — | — | — | 672.2x |
| 29 | Granite 4.0 1B Speech | IBM | 53.9 | 2 of 8 | 5.1% | — | — | — | — | — | — | 661.2x |
| 30 | Nemotron Speech Streaming EN 0.6B | NVIDIA | 53.4 | 2 of 8 | 5.7% | — | — | — | — | — | — | 1071.4x |
| 31 | Pulse (partial coverage) | Smallest AI | 52.9 | 1 of 8 | 5.4% | — | — | — | — | — | — | — |
| 32 | Parakeet CTC 1.1B | NVIDIA | 52.2 | 3 of 8 | 6.5% | 12.9% | — | — | — | — | — | 5014.5x |
| 33 | Qwen3 ASR 1.7B | Qwen | 52.1 | 2 of 8 | 5.0% | — | — | — | — | — | — | 394.1x |
| 34 | Phi 4 Multimodal Instruct | Microsoft | 52.1 | 7 of 8 | 5.4% | — | 5.0% | 4.1% | 3.8% | 4.3% | 5.2% | 162.7x |
| 35 | Canary 180M Flash | NVIDIA | 51.3 | 2 of 8 | 6.3% | — | — | — | — | — | — | 2484.2x |
| 36 | Qwen3 ASR 1.7B HF | Qwen | 51.0 | 7 of 8 | 5.0% | — | 5.7% | 4.0% | 3.8% | 5.4% | 6.3% | 796.2x |
| 37 | Whisper Large V3 | OpenAI | 50.9 | 8 of 8 | 6.5% | 11.2% | 6.2% | 4.0% | 3.4% | 4.5% | 4.9% | 462.2x |
| 38 | MOSS Transcribe Preview 2B | OpenMOSS | 50.3 | 2 of 8 | 4.7% | — | — | — | — | — | — | 148.5x |
| 39 | MOSS Transcribe Diarize | OpenMOSS | 50.0 | 2 of 8 | 5.2% | — | — | — | — | — | — | 382.4x |
| 40 | Higgs Audio v3 STT | Boson AI | 49.9 | 2 of 8 | 4.6% | — | — | — | — | — | — | 110.1x |
| 41 | Parakeet CTC 0.6B | NVIDIA | 49.9 | 3 of 8 | 6.7% | 13.7% | — | — | — | — | — | 5883.9x |
| 42 | Hojo ASR v1 | Hojo AI | 49.6 | 2 of 8 | 4.5% | — | — | — | — | — | — | 73.8x |
| 43 | Moonshine Streaming Small | Useful Sensors | 49.3 | 2 of 8 | 6.8% | — | — | — | — | — | — | 3206.0x |
| 44 | Distil Large v3.5 | Distil Whisper | 49.2 | 3 of 8 | 6.1% | 11.7% | — | — | — | — | — | 874.0x |
| 45 | Parakeet Tdt 1.1B | NVIDIA | 49.0 | 3 of 8 | 6.2% | 15.8% | — | — | — | — | — | 4528.7x |
| 46 | Canary 1B | NVIDIA | 48.9 | 2 of 8 | 5.8% | — | — | — | — | — | — | 765.6x |
| 47 | Granite Speech 3.3 2B | IBM | 48.6 | 2 of 8 | 5.5% | — | — | — | — | — | — | 507.1x |
| 48 | Higgs Audio v3 8B STT v2 | Boson AI | 48.5 | 2 of 8 | 4.7% | — | — | — | — | — | — | 139.1x |
| 49 | Niagara 38M Batch.en | Applied Brain Research | 47.8 | 2 of 8 | 8.3% | — | — | — | — | — | — | 4048.7x |
| 50 | OmniASR LLM 7B v2 | Meta | 47.7 | 7 of 8 | 6.9% | — | 5.3% | 4.4% | 3.4% | 3.6% | 5.2% | 142.9x |
| 51 | Parakeet Rnnt 0.6B | NVIDIA | 47.7 | 3 of 8 | 6.7% | 14.5% | — | — | — | — | — | 5406.7x |
| 52 | Moonshine Streaming Tiny | Useful Sensors | 47.6 | 2 of 8 | 11.2% | — | — | — | — | — | — | 4375.2x |
| 53 | Granite Speech 3.3 8B | IBM | 47.4 | 2 of 8 | 5.3% | — | — | — | — | — | — | 263.7x |
| 54 | Niagara 19M Batch.en | Applied Brain Research | 46.7 | 2 of 8 | 9.9% | — | — | — | — | — | — | 3735.5x |
| 55 | Qwen3 ASR 0.6B | Qwen | 46.7 | 2 of 8 | 5.6% | — | — | — | — | — | — | 438.7x |
| 56 | Parakeet Rnnt 1.1B | NVIDIA | 46.4 | 3 of 8 | 6.4% | 17.1% | — | — | — | — | — | 4121.8x |
| 57 | Zipformer Cr CTC Transducer XL (290M) | SoundsGood AI | 46.1 | 2 of 8 | 5.2% | — | — | — | — | — | — | 159.8x |
| 58 | Moonshine Tiny | Useful Sensors | 45.6 | 2 of 8 | 11.4% | — | — | — | — | — | — | 3733.1x |
| 59 | Chirp (partial coverage) | Google | 45.5 | 1 of 8 | — | 13.1% | — | — | — | — | — | — |
| 60 | Whisper Large V3 Turbo | OpenAI | 44.9 | 8 of 8 | 7.0% | 11.0% | 6.7% | 6.1% | 3.9% | 5.4% | 4.6% | 782.6x |
| 61 | Hubert Large Ls960 Ft | Meta | 44.1 | 2 of 8 | 13.4% | — | — | — | — | — | — | 3024.0x |
| 62 | Wav2vec2 Large 960h Lv60 Self | Meta | 44.0 | 2 of 8 | 11.8% | — | — | — | — | — | — | 3010.0x |
| 63 | STT EN Conformer Transducer Small (partial coverage) | NVIDIA | 42.8 | 1 of 8 | — | 13.7% | — | — | — | — | — | — |
| 64 | Owsm CTC v4 1B | ESPnet | 42.6 | 2 of 8 | 6.7% | — | — | — | — | — | — | 764.7x |
| 65 | Voxtral Mini 3B 2507 | Mistral | 42.5 | 7 of 8 | 6.0% | — | 6.0% | 4.5% | 4.2% | 5.1% | 5.1% | 179.7x |
| 66 | Owsm CTC v3.1 1B | ESPnet | 42.0 | 2 of 8 | 7.3% | — | — | — | — | — | — | 816.2x |
| 67 | GLM ASR Nano 2512 | Z.ai | 41.8 | 2 of 8 | 6.0% | — | — | — | — | — | — | 333.0x |
| 68 | STT EN Conformer CTC Large (partial coverage) | NVIDIA | 41.4 | 1 of 8 | — | 14.4% | — | — | — | — | — | — |
| 69 | Mms 1B All | Meta | 41.3 | 2 of 8 | 13.5% | — | — | — | — | — | — | 1954.0x |
| 70 | Audio8 ASR 0.1B | AutoArk AI | 40.1 | 2 of 8 | 7.0% | — | — | — | — | — | — | 709.3x |
| 71 | STT 2.6B EN | Kyutai | 39.6 | 2 of 8 | 5.7% | — | — | — | — | — | — | 133.2x |
| 72 | Owsm CTC v3.2 Ft 1B | ESPnet | 39.2 | 2 of 8 | 7.3% | — | — | — | — | — | — | 692.3x |
| 73 | OmniASR LLM 3B v2 | Meta | 39.1 | 5 of 8 | — | — | 6.9% | 5.6% | 4.0% | 5.0% | 7.1% | — |
| 74 | Lite Whisper Large v3 Acc | Efficient Speech | 39.1 | 2 of 8 | 6.3% | — | — | — | — | — | — | 203.9x |
| 75 | Zipformer Transducer XL (290M) | SoundsGood AI | 37.3 | 2 of 8 | 6.2% | — | — | — | — | — | — | 141.2x |
| 76 | Crisperwhisper | Nyra Health | 36.8 | 2 of 8 | 5.8% | — | — | — | — | — | — | 33.3x |
| 77 | STT EN Conformer CTC Small (partial coverage) | NVIDIA | 36.1 | 1 of 8 | — | 18.7% | — | — | — | — | — | — |
| 78 | OmniASR LLM 1B v2 | Meta | 35.2 | 5 of 8 | — | — | 7.0% | 6.3% | 4.3% | 4.8% | 7.3% | — |
| 79 | OmniASR CTC 7B v2 | Meta | 35.0 | 7 of 8 | 9.0% | — | 7.6% | 6.0% | 4.4% | 4.7% | 6.5% | 527.6x |
| 80 | STT EN Fastconformer CTC Large (partial coverage) | NVIDIA | 34.8 | 1 of 8 | — | 21.5% | — | — | — | — | — | — |
| 81 | STT EN Conformer Transducer Large (partial coverage) | NVIDIA | 33.5 | 1 of 8 | — | 21.9% | — | — | — | — | — | — |
| 82 | OmniASR CTC 3B v2 | Meta | 33.2 | 5 of 8 | — | — | 8.1% | 6.2% | 4.5% | 5.1% | 6.5% | — |
| 83 | STT EN Fastconformer Transducer Large (partial coverage) | NVIDIA | 32.1 | 1 of 8 | — | 22.5% | — | — | — | — | — | — |
| 84 | Voxtral Mini 4B Realtime 2602 | Mistral AI | 30.4 | 7 of 8 | 6.4% | — | 7.8% | 6.3% | 4.2% | 6.4% | 6.2% | 105.1x |
| 85 | Qwen3 ASR 0.6B HF | Qwen | 28.7 | 7 of 8 | 5.6% | — | 8.8% | 6.7% | 5.8% | 9.1% | 10.1% | 730.2x |
| 86 | ASR Conformer Largescaleasr | SpeechBrain | 27.6 | 2 of 8 | 7.6% | — | — | — | — | — | — | 72.4x |
| 87 | Nemotron 3.5 ASR Streaming 0.6B | NVIDIA | 26.5 | 7 of 8 | 7.9% | — | 9.4% | 9.0% | 5.4% | 8.7% | 8.3% | 1489.6x |
| 88 | OmniASR CTC 1B v2 | Meta | 23.0 | 5 of 8 | — | — | 10.3% | 8.4% | 5.6% | 6.4% | 9.0% | — |
| 89 | OmniASR LLM 300M v2 | Meta | 20.5 | 5 of 8 | — | — | 10.0% | 9.1% | 5.8% | 6.9% | 9.5% | — |
| 90 | Vibevoice ASR HF | Microsoft | 20.4 | 7 of 8 | 6.3% | — | 14.0% | 14.2% | 6.7% | 12.3% | 7.8% | 221.2x |
| 91 | OmniASR CTC 300M v2 | Meta | 14.3 | 5 of 8 | — | — | 18.6% | 16.5% | 9.9% | 11.5% | 14.4% | — |

## How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 2 of the 8 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

## Data sources

- [Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard): Word error rates from the Hugging Face Open ASR Leaderboard.
