Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

EN | 中文

CuteTTS logo CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

CuteTTS Hugging Face model CuteTTS-distill Hugging Face model paper

  • A lightweight (~230M-parameter) continuous autoregressive TTS model that runs efficiently on GPUs, CPUs, and Apple silicon.
  • Ultra-low latency: ~40 ms to the first audio chunk and a throughput of ~9× real time on an NVIDIA RTX 4090.
  • Excellent speech quality and voice cloning performance.
  • Web demo, Python API, and CLI.
  • Multilingual support: English, Chinese, French, German, and Spanish.

CuteTTS architecture


CuteTTS performance

Install

Using conda

conda create -n cutetts python=3.12 -y
conda activate cutetts
pip install torch==2.5.1 torchaudio==2.5.1  # For NVIDIA GPUs with CUDA 12.1, append: --index-url https://download.pytorch.org/whl/cu121
pip install -e .

Download weights

Download CuteTTS or CuteTTS-distill into ./model:

mkdir -p ./model
hf download OPPOer/CuteTTS --local-dir ./model/CuteTTS
hf download OPPOer/CuteTTS-distill --local-dir ./model/CuteTTS-distill

Web demo

cutetts-demo --model-dir ./model --device auto --host 127.0.0.1 --port 7860

Python

from cutetts import CuteTTS
import soundfile as sf

model = CuteTTS.from_pretrained("/path/to/CuteTTS", device="auto")
result = model.generate("The voice generated by this model sounds amazing!", mode="tts")
sf.write("tts.wav", result.waveform.squeeze(0).numpy(), result.sample_rate)

clone = model.generate(
    "The voice generated by this model sounds amazing!",
    mode="voice_clone",
    reference_audio="assets/default_reference.wav",
)
sf.write("clone.wav", clone.waveform.squeeze(0).numpy(), clone.sample_rate)

Streaming Python API

generate_stream() yields decoded CPU float32 PCM chunks as soon as they are ready:

stream = model.generate_stream(
    "The voice generated by this model sounds amazing!",
    mode="voice_clone",
    reference_audio="assets/default_reference.wav",
)

with sf.SoundFile(
    "stream.wav",
    mode="w",
    samplerate=model.sample_rate,
    channels=1,
    subtype="PCM_16",
) as output:
    for chunk in stream:
        output.write(chunk.waveform.squeeze(0).numpy())

Command line

cutetts --model-dir /path/to/CuteTTS --mode tts --text "The voice generated by this model sounds amazing!" --output tts.wav

cutetts --model-dir /path/to/CuteTTS-distill \
  --mode voice_clone --reference-audio assets/default_reference.wav \
  --text "The voice generated by this model sounds amazing!" --output clone.wav

Zero-shot voice-cloning performance

Model Params. LibriSpeech test-clean WER (%) ↓ LibriSpeech test-clean SIM ↑ Seed-TTS EN WER (%) ↓ Seed-TTS EN SIM ↑ Seed-TTS ZH WER (%) ↓ Seed-TTS ZH SIM ↑
MOSS‑TTS 8B 1.98 67.7 1.84 70.9 1.37 77.0
Qwen3‑TTS 1.7B 2.35 70.3 1.66 71.4 0.91 77.0
FireRedTTS‑2 1.5B 4.32 64.2 1.95 66.5 1.14 73.6
MOSS‑TTS‑Nano 0.1B 4.10 48.4 4.62 49.9 3.13 64.3
F5‑TTS 0.3B 2.42 66.0 1.83 67.0 1.56 76.0
ZipVoice 0.1B 2.05 67.4 1.70 69.7 1.40 75.1
IndexTTS2 1.5B 2.47 70.0 2.22 70.6 1.02 76.5
CosyVoice 3 0.5B 1.99 69.7 2.02 71.8 1.16 78.0
VoxCPM2 2B 3.01 74.0 1.84 75.3 0.97 79.5
VibeVoice 1.5B – – 3.04 68.9 1.16 74.4
DiTAR 0.6B 2.39 67.0 1.69 73.5 1.02 75.3
VibeVoice‑Realtime 0.5B 2.00 69.5 2.05 63.3 – –
Pocket TTS 0.1B 1.59 49.1 1.63 50.7 – –
CuteTTS 0.2B 2.16 78.9 2.04 76.5 1.41 77.8
CuteTTS‑distill 0.2B 2.41 76.8 2.03 74.2 1.47 75.6

Acknowledgements

License

CuteTTS is released under the Apache License 2.0. Third-party portions retain their respective copyright notices and licenses as described in NOTICE.

About

CuteTTS: a lightweight continuous autoregressive TTS model that runs efficiently on GPUs, CPUs, and Apple silicon.

Resources

Stars

80 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages