← VoiceHub ArenaDownload investigation ↓
ADDITIONAL QUALITY AUDIT · 15 SEPTEMBER 2026

High error rates.
Checked against the evidence.

132 diagnostic records scored and verified; 96 newly generated recordings. All 35,904 original records were checked. IDs, target texts and every published WER/CER calculation matched. We then examined six additional models above 3% WER and verified all 6,528 source WAVs.

ModelFull archived WERWrong wordsMissing wordsExtra wordsFinding
Vui Abraham 100M12.15%528182741Short-text and minimum-duration sensitivity; residual quality failures.
ConversationTTS · ckpt18.72%21629796Confirmed duration-budget bug repaired; strong speaker sensitivity.
Bark Small7.74%456128340Independent waveforms match exactly; ASR amplifies some failures.
VITS · MMS English6.23%6065880No precision-related WER improvement; independent implementation matched.
OpenVoice V23.68%3723533Conversion stage degrades intelligibility; publisher source embedding partly helps.
VoxCPM23.35%9914287Tail continuation/repetition; higher CFG did not consistently help.

Which errors dominate?

Download SVG ↓
Stacked word-error components for six models, using all 1088 original texts per model

Vui Abraham 100M

This is Vui Abraham 100M (vui-abraham-100m.pt), not a newer Vui checkpoint. Across the full corpus, texts below 40 characters have 43.49% WER versus 6.67% at 80–119 characters. The native and pinned publisher source both suppress EOS for roughly four seconds. In six deliberately difficult cases, lowering this floor to one second reduced diagnostic WER from 80.52% to 58.44% by reducing inserted words from 22 to 5; substitutions and deletions did not improve. The original baseline reproduced its archived score. This is evidence of an unsuitable stopping constraint for some short texts, not proof that the remaining failures are fixed. The one-second change was an in-memory experiment, not a deployed default.

Publisher reference

ConversationTTS · ckpt1

The actual model is AudioFoundation/SpeechFoundation ckpt1. Its generator budgeted 40 ms per audio frame, while the actual Mimi codec produced 80 ms per frame (20 frames decoded to 38,400 samples at 24 kHz). This allowed a requested duration to be exceeded by up to two times. The generator now derives its budget from the codec frame rate, and wrapper validation uses 80 ms. Three non-terminating-model regression cases failed before and passed after the repair. Only two archived full-corpus recordings exceeded 30 seconds, so this bug alone cannot explain the full 8.72% WER. Separately, changing only speaker=0 to speaker=1 reduced WER on the six selected cases from 172.22% to 6.94%, and extra words from 119 to 1. A 41.12-second failure became 3.28 seconds. This is a promising configuration finding, not a corrected full-split score or proof that speaker 0 is invalid. The publisher describes training labels [1] and [2], while the upstream API defaults to 0. Both diagnostic variants were generated before the duration patch; all original recordings remain unchanged.

Publisher reference

Bark Small

The archived model is Bark Small, suno/bark-small, without a fixed voice/history preset. All six independent Transformers FP32 outputs were sample-for-sample identical to the native regenerated waveforms, and native baselines matched the archived waveforms. The same diagnostic WER of 122.37% is therefore not explained by a discrepancy between these tested implementations. Turning VAD on for the same archived audio reduced selected-case WER to 48.68%; one four-word target triggered a long repeated Whisper transcript. Other cases still continued beyond the target. Recognition amplification contributes to extreme scores, while remaining synthesis/conditioning failures are unresolved. No silent VAD change or preset substitution was made to the full benchmark.

Publisher reference

VITS · MMS English

The checkpoint is facebook/mms-tts-eng. Native and Transformers tokenizers matched on all 1,088 texts. On six selected CPU FP32 cases, independent waveforms had identical lengths and near-unit correlation; the largest absolute sample difference was 2.95043e-5. Their diagnostic WER and CER also matched. On CUDA, the original native dtype was FP16: regenerating in FP32 left diagnostic WER unchanged at 22.97%. CPU and CUDA use different random-number streams, so the CPU/CUDA score difference is not a precision comparison. These checks do not support an FP16 failure or a material native-port discrepancy in the tested path. Errors such as names and word pronunciations remain; the six cases cannot rule out every possible defect.

Publisher reference

OpenVoice V2

This is OpenVoice V2 with the Melo English base, default EN-US speaker (id 0), and the fixed Emily target reference. We saved the exact Melo waveform entering each conversion. Across the six selected cases, Melo alone had 2.67% WER, compared with 26.67% after conversion. The converted waveform exactly reproduced the archived output. The native default estimates a source speaker embedding separately from each short generated clip, whereas the publisher V2 example uses a released embedding for the known Melo speaker. Substituting only the released en-us.pth source embedding reduced diagnostic WER to 14.67%. One selected failure improved substantially, but two others remained poor. This establishes conversion/conditioning sensitivity and a partial workflow difference, not a complete fix or independent numerical converter-parity proof. The target reference and its processing remain possible factors.

Publisher reference

VoxCPM2

The archived checkpoint is VoxCPM2, using BF16 and CFG=2. Insertions account for 287 of 400 word edits in the full corpus. The three highest-edit cases initially transcribe the complete target, then add unrelated or repeated text; VAD did not eliminate these failures. Increasing CFG from 2 to 3 reduced some tails and raised exact matches from three to four of the six cases, but a severe repeated-transcription failure increased their combined diagnostic WER from 39.76% to 134.94%. Therefore CFG=3 is not a validated fix and was not adopted. Generation termination and recognition amplification remain under investigation; this sensitivity test is not independent publisher-parity validation.

Publisher reference

Controlled diagnostic comparisons

Three highest-edit-count cases, first two records and longest text per model. Seed 42, pinned Whisper-large-v3, English, FP16, no VAD. These percentages describe selected difficult cases only.

ModelVariantRecordsDiagnostic WERDiagnostic CER
Vui Abraham 100Marchived680.52%57.51%
Vui Abraham 100Mbaseline680.52%57.51%
Vui Abraham 100Mminimum_1_second658.44%43.43%
ConversationTTS · ckpt1archived6172.22%147.70%
ConversationTTS · ckpt1baseline6172.22%147.70%
ConversationTTS · ckpt1speaker_166.94%3.57%
Bark Smallarchived6122.37%92.36%
Bark Smallbaseline6122.37%92.36%
Bark Smalltransformers6122.37%92.36%
VITS · MMS Englisharchived622.97%9.73%
VITS · MMS Englishbaseline622.97%9.73%
VITS · MMS Englishfp32622.97%9.49%
VITS · MMS Englishnative_cpu_fp32620.27%9.49%
VITS · MMS Englishtransformers_cpu_fp32620.27%9.49%
OpenVoice V2archived626.67%12.97%
OpenVoice V2baseline626.67%12.97%
OpenVoice V2captured_conversion626.67%12.97%
OpenVoice V2melo_before_conversion62.67%0.50%
OpenVoice V2publisher_source_embedding614.67%7.73%
VoxCPM2archived639.76%42.27%
VoxCPM2baseline639.76%42.27%
VoxCPM2cfg_36134.94%56.14%

Hear selected diagnostic examples

These examples illustrate failure mechanisms; the six-case and full-corpus tables above provide the denominators.

ConversationTTS · ckpt1

Target: The doctor cried after his birth.

Original speaker 0

ASR: The doctor cried after his birth, and they, ant of his fate, crap, and it's, and pename stover, right? He had various, vil tarroid on the very point, and, and telt, they been in wrecks, and fight Dr. Water, so this adjuvant and this dartig, and Dr. Reporter.

Speaker 1

ASR: The doctor cried after his birth.

OpenVoice V2

Target: Kanwal is said to mean "snakes indeed" in a local Aboriginal language.

Melo before conversion

ASR: Kenwal is said to mean snakes indeed in a local Aboriginal language.

Original conversion

ASR: Can we all accept the means of X and T in a local Aboriginal language?

Released source embedding

ASR: Kenwell is said to mean snakes indeed in a local Aboriginal language.

Inspect the evidence

Complete report and sources · All 33 model record checks · 6,528 WAV checks · Independent VITS comparison · Diagnostic transcripts and metrics · Audio and metric verification · Duration regression proof · All diagnostic audio

All percentages in the main six-model table describe the archived full 1,088-text benchmark. Selected diagnostic results do not replace them. WER may exceed 100% when insertions exceed the number of target words. VAD scores are sensitivity checks, not a new scoring protocol. No human transcription or perceptual-quality rating was collected in this audit; ASR output alone cannot establish precisely which repetitions are audible. The benchmark uses fixed provider voices/references rather than the official per-prompt speaker-identity protocol. Broader resynthesis is required before publishing corrected full-split scores for any newly proposed configuration.