---
title: "Best AI for text to speech"
description: "Speech models ranked on TTS Arena V2 Elo from blind pairwise listening votes, with entries still marked preliminary upstream left out."
url: "https://aldena.ai/best-ai-for-text-to-speech"
---

# Best AI for text to speech

Blind pairwise listening votes. Models still marked preliminary upstream are left out until their rating settles.

Aldena runs these models inside your team rooms. [See what each one costs](https://aldena.ai/models).

| rank | model | vendor | score | coverage (benchmarks) | TTS Arena V2 (elo) |
| --- | --- | --- | --- | --- | --- |
| 1 | Luna TTS | Vui Labs | 66.7 | 1 of 1 | 1574 |
| 2 | Luck Dolphin | Luck Dolphin | 65.7 | 1 of 1 | 1561 |
| 3 | CastleFlow v1.0 | Async | 64.6 | 1 of 1 | 1560 |
| 4 | Inworld TTS (max) | Inworld | 63.6 | 1 of 1 | 1558 |
| 5 | Luck Dolphin Turbo | Luck Dolphin | 62.6 | 1 of 1 | 1550 |
| 6 | Papla P1 | Papla | 61.6 | 1 of 1 | 1549 |
| 7 | Inworld TTS | Inworld | 60.6 | 1 of 1 | 1543 |
| 8 | Speech 2.8 HD | MiniMax | 59.6 | 1 of 1 | 1536 |
| 9 | Lightning v3.1 Pro | Lightning | 58.6 | 1 of 1 | 1530 |
| 10 | Hume Octave | Hume | 57.6 | 1 of 1 | 1527 |
| 11 | Gradium TTS | Gradium | 56.6 | 1 of 1 | 1525 |
| 12 | Speech 2.8 Turbo | MiniMax | 55.6 | 1 of 1 | 1523 |
| 13 | OpenAudio S2 | Lanternfish | 54.5 | 1 of 1 | 1522 |
| 14 | MiniMax Speech 2.6 Turbo | MiniMax | 53.5 | 1 of 1 | 1520 |
| 15 | Cartesia Sonic 2 | Cartesia | 52.5 | 1 of 1 | 1517 |
| 16 | MiniMax Speech 2.6 HD | MiniMax | 51.0 | 1 of 1 | 1514 |
| 17 | Typecast SSFM 3.0 | Typecast | 51.0 | 1 of 1 | 1514 |
| 18 | MiniMax Speech 02 HD | MiniMax | 49.5 | 1 of 1 | 1513 |
| 19 | OpenAudio S1 | Lanternfish | 48.5 | 1 of 1 | 1512 |
| 20 | Eleven Turbo v2.5 | ElevenLabs | 46.5 | 1 of 1 | 1511 |
| 21 | Hithink Speech 2.6 | Hithink | 46.5 | 1 of 1 | 1511 |
| 22 | Inworld TTS 1.5 MAX | Inworld | 46.5 | 1 of 1 | 1511 |
| 23 | Eleven Flash v2.5 | ElevenLabs | 43.4 | 1 of 1 | 1508 |
| 24 | Eleven Multilingual v2 | ElevenLabs | 43.4 | 1 of 1 | 1508 |
| 25 | Eleven v3 | ElevenLabs | 43.4 | 1 of 1 | 1508 |
| 26 | MiniMax Speech 02 Turbo | MiniMax | 41.4 | 1 of 1 | 1504 |
| 27 | Voice.ai Text to Speech V1 | Voice.ai | 40.4 | 1 of 1 | 1498 |
| 28 | Chatterbox | Resemble AI | 39.4 | 1 of 1 | 1480 |
| 29 | Kokoro v1.0 | Hexgrad | 38.4 | 1 of 1 | 1478 |
| 30 | Magpie Research Preview | Magpie | 37.4 | 1 of 1 | 1474 |
| 31 | NeuTTS Max | Neuphonic | 36.4 | 1 of 1 | 1440 |
| 32 | Maya 1 | Maya Research | 35.4 | 1 of 1 | 1411 |
| 33 | Magpie Multilingual | Magpie | 34.3 | 1 of 1 | 1406 |
| 34 | Wordcab TTS | Wordcab | 33.3 | 1 of 1 | 1361 |

## How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 1 of the 1 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

## Data sources

- [TTS Arena V2](https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2): Elo ratings from TTS Arena V2.
