Xiaohongshu Tech REDtech
Aug 10, 2022 · Artificial Intelligence
Multi-Stage Multi-Codebook VQ-VAE for High-Performance Neural Text-to-Speech (MSMC‑TTS)
The MSMC‑TTS system, a multi‑stage multi‑codebook VQ‑VAE based neural text‑to‑speech solution, delivers near‑human audio quality (MOS 4.41) with a compact 3.12 MB acoustic model, substantially surpassing Mel‑Spectrogram FastSpeech baselines in naturalness and efficiency.
Compact RepresentationMulti-Stage ModelingNeural TTS
0 likes · 10 min read