VoiceHub Arena
COMPLETED BENCHMARK / SEPTEMBER 2026

One dataset.
Every voice, measured.

Compare 33 text-to-speech models on the complete English Seed-TTS-Eval split. Every score has a transcript. Every transcript has a recording.

EVALUATION PROTOCOLSeed-TTS-Eval · English

Whisper-large-v3 transcription
Fixed provider voices & references
Same 1,088 texts for every model

Read the methodology ↓
Models evaluated33/ 33
Texts per model1,088
Scored recordings35,904
Evaluation hardwareA100 40 GB
A CLOSER LOOK

The voices, side by side.

Choose a metric. Compare the bars. Inspect the full table below.

Download plots ↓

Same texts. Same evaluator.
1,088 samples for every model.

Word error rate

Loading chart…

Choose models
View the values behind this chart
ModelWER95% confidence interval

The full comparison

Corpus error rates. Lower WER, CER and RTF are better.

Loading results…

Select a model name to listen to its samples. WER/CER intervals are 95% prompt bootstrap intervals; small rank differences may not be meaningful. Review marks a confirmed or unresolved quality issue. CosyVoice’s archived score has no rank.

HOW TO READ THESE RESULTS

One experiment. A transparent record.

All recordings & records ↗

Matched evaluation

All 33 models synthesize the same 1,088 English texts with seed 42 and one repeat. WER and CER aggregate edit counts across the corpus, rather than averaging shard percentages. Recognition uses pinned Whisper-large-v3 and whisper_english normalization.

What the scores mean

WER measures word edits; CER measures character edits. Both estimate intelligibility through ASR. They do not measure naturalness or speaker similarity. Fixed provider voices and references mean this is not the official zero-shot speaker-identity/SIM protocol.

Keep the difficult cases

CosyVoice has a confirmed vocoder defect in this archived run. Dia and Llasa completed inference but have high transcript error rates whose causes are still under review. Their outputs are included. Empty ASR text alone does not prove silence. Successful inference is not a guarantee of good speech.

Metric definitions & reproducibility

MER / WIL / WIP: match error rate, word information lost and word information preserved. Lower MER/WIL and higher WIP are better. Exact match: fraction of normalized target/transcript pairs that match exactly.

RTF: summed synthesis seconds / summed audio seconds; below 1 is faster than real time. Latency p50/p95: synthesis time percentiles. Silence/clipping: measured signal diagnostics, not subjective quality ratings.

ASR: Systran/faster-whisper-large-v3, revision edaa852ec7e145841d8ffdb056a99866b5f0a478; CUDA FP16, English, beam 5, temperature 0, no VAD or previous-text conditioning. 95% intervals: 1,000 prompt-cluster bootstrap resamples, seed 42.

Coverage verification · Audio hash verification · Original Seed-TTS-Eval