開放式TTS排行榜:多語言語音合成與聲音克隆的可擴展評測

2026年9月30日 00:00
站內 AI 整理稿

Back to Articles Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning Published September 30, 2026 Update on GitHub Upvote 2 Eric Bezzam bezzam Follow Steven Zheng Steveeeeeeen Follow Eustache Le Bihan eustlb Follow mrfakename mrfakename Follow guest The pace of open-source text-to-speech (TTS) model releases has been incredible.

On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀 Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized.The gold standard is human preference scores such as MOS or MUSHRA (more on metrics).

To this end, several arena-based leaderboards have established themselves as useful reference points for the community: TTS Arena v2 Artificial Analysis Voice Arena These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other.

After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model (see Voice Arena methodology).While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases.

This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena.

This likely reflects practical factors: adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors.

Another limitation with arena-style evaluation is voter consistency: no arena can ensure that the same voters with the same criteria of “better” can consistently evaluate models over time.

Even the preferences of a single person change over time (“A man cannot step into the same river twice” as famously said by Heraclitus).

To this end, we've built the Open TTS Leaderboard, which uses objective metrics to evaluate models on complementary aspects of performance: Intelligibility: word/character error rate (WER and CER) between the prompt and the generated audio's transcript, using Qwen3 ASR (top ranking open-source model on the Open ASR Leaderboard).

Speed: inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU.

Speaker similarity by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.

By relying on objective metrics evaluating a model drops from a couple weeks (for collecting votes) to a couple hours ⚡ Importantly, the Open TTS Leaderboard does not replace human preference ranking.

ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation.Neither directly measures naturalness, expressiveness, or listener preference.Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.

Our intention with this leaderboard is for it to be shaped by the community; we want to hear your feedback so the evaluations stay relevant and insightful.The next few sections give an overview of main features of the Open TTS Leaderboard.

Multilingual + voice cloning evaluation From the default view of the leaderboard, models are ranked by macro-average WER on the English splits of Seed TTS Eval (paper) and CV3 Eval (zero shot) (paper).

hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits, while the Pareto plots visualize which models strike a good balance between WER, batched inference (RTFx), and size.

English performance doesn't necessarily translate to other languages.Multiple languages can be toggled to rank models on multilingual performance.Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot).

Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models.

By toggling “Voice cloning”, the models that support this functionality (on the selected languages) can be compared.Moreover, a SIM column for speaker similarity now appears in the table, as well as two more Pareto plots for visualizing the tradeoff between SIM, batched inference, and size.

The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improve under voice cloning, namely when a reference audio is provided.Compare and vote on TTS outputs Numbers only tell part of the story, and as mentioned earlier human preference is the ultimate decider.

From the “Listen” tab, you can compare the generated outputs that are behind the metrics, to find which model(s) you prefer!Pick the language/dataset you're interested in, whether you want to compare voice cloning, and optionally pick the models or listen to outputs from a random selection.

The “Listen” tab fills an important gap in existing TTS leaderboards: a space to explore model outputs of various models.You can even give feedback on the generated outputs.As we collect more votes from the community, we may include this data on the leaderboard.So vote!

But please login with your HF account to help us weed out spam/bots.Streaming performance The “Streaming” tab compares the streaming capabilities.Models are ranked by TTFA (time-to-first-audio), which quantifies how long a user waits after probing a model in order to obtain audio that can be played.

This is important for voice agents and other interactive apps.For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives.For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier.

Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice.We drop the first 3 runs as warm-up and report the median TTFA across the rest.The default view compares performance on an H200 GPU.

Results are also available for CPU for a small (but growing) set of models!kyutai/pocket-tts is a great model for streaming on both GPU and CPU!

Conclusion The goal of the Open TTS Leaderboard is not only to keep up with the incredible pace of TTS model releases, but to be shaped by the community; we want to hear your feedback so the evaluations stay relevant and insightful.Let us know which datasets, models, and metrics you want to see!

For now, we've focused on: Open-source models, to put forward many great models that have been neglected by arena-style evaluations.Multilingual, since English performance is not a suitable proxy for other languages.

We will soon open-source the evaluation scripts, much like the Open ASR Leaderboard repo, so that you can directly provide your feedback and suggestions via GitHub Issues and PRs!

Let's shape TTS evaluations together 🤗 Models mentioned in this article 10 Spaces mentioned in this article 3 Papers mentioned in this article 2 More Articles from our Blog audiospeechleaderboard The Open ASR Leaderboard Adds Its First Global South Language +6 59 August 28, 2026 audiospeechbenchmark Introducing Real World VoiceEQ: Measuring the human quality of voice AI +9 33 July 15, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 2 Models mentioned in this article 10 Spaces mentioned in this article 3 Papers mentioned in this article 2

Related

相關文章

Hugging Face Blog自然語言處理

像物理學家一樣剪枝大型語言模型:將區塊移除視為伊辛最佳化問題

讓大型語言模型加速最簡單粗暴的方法之一,就是刪除整個轉換器區塊。由於模型實際變短,區塊移除(亦稱深度剪枝)不僅節省記憶體,還能帶來可預測的推論加速,並能與量化、低秩壓縮等其他技術乾淨地疊加。難點在於決定要剪掉哪些區塊。剪錯區塊,模型就會。

1 週前

人格幾何與對齊:讓模型的內部結構與人類認知結構對齊的可能性

當我們看到模型內部的人格結構與人類高度一致時,一個自然的問題是,為什麼這種結構會在模型中出現。研究團隊給出了一個非常有說服力的解釋。語言中的人格詞彙結構高度穩定,人類在長期的社會互動中形成了共同的隱性人格結構,而語言模型從海量文本中學習語言時,自然繼承了這種結構。

1 週前

微軟及 OpenAI 內部文件曝光:高管曾警告 AI 可能對新聞機構產生“毀滅性影響”

微軟與 OpenAI 的內部文件近日曝光,揭露高層曾對 AI 技術可能對新聞產業帶來的衝擊提出嚴厲警告。根據文件內容,微軟應用科學部門主管 Brent Hecht 直言,將新聞內容用於 AI 模型訓練的行為,等同於「前所未有的大規模盜竊」。這份內部記錄還顯示,微軟掌握的數據指出,部分正在對 AI 公司提起版權訴訟的新聞機構,其網站點擊率已大幅下滑超過 80%。 這些內部文件進一步凸顯了 AI 訓練過程中使用新聞素材所引發的法律與道德爭議。

1 週前

被英偉達點名的杭州團隊,補上了AI for Science的「最後一公里」

一家來自杭州的團隊,近期獲得英偉達的公開關注,其技術被認為補齊了AI for Science(科學智慧)領域中「最後一公里」的關鍵環節。該團隊開發的系統,能讓科學研究者直接透過對話式互動,從最初的研究想法快速產出具體結果,大幅縮短了過去需要大量編碼與繁複流程的轉化路徑。 這項技術的核心在於將自然語言理解與科學運算流程深度整合。

2 週前