CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
- A lightweight (~230M-parameter) continuous autoregressive TTS model that runs efficiently on GPUs, CPUs, and Apple silicon.
- Ultra-low latency: ~40 ms to the first audio chunk and a throughput of ~9× real time on an NVIDIA RTX 4090.
- Excellent speech quality and voice cloning performance.
- Web demo, Python API, and CLI.
- Multilingual support: English, Chinese, French, German, and Spanish.
Using conda
conda create -n cutetts python=3.12 -y
conda activate cutetts
pip install torch==2.5.1 torchaudio==2.5.1 # For NVIDIA GPUs with CUDA 12.1, append: --index-url https://download.pytorch.org/whl/cu121
pip install -e .Download weights
Download CuteTTS or CuteTTS-distill into ./model:
mkdir -p ./model
hf download OPPOer/CuteTTS --local-dir ./model/CuteTTS
hf download OPPOer/CuteTTS-distill --local-dir ./model/CuteTTS-distillcutetts-demo --model-dir ./model --device auto --host 127.0.0.1 --port 7860from cutetts import CuteTTS
import soundfile as sf
model = CuteTTS.from_pretrained("/path/to/CuteTTS", device="auto")
result = model.generate("The voice generated by this model sounds amazing!", mode="tts")
sf.write("tts.wav", result.waveform.squeeze(0).numpy(), result.sample_rate)
clone = model.generate(
"The voice generated by this model sounds amazing!",
mode="voice_clone",
reference_audio="assets/default_reference.wav",
)
sf.write("clone.wav", clone.waveform.squeeze(0).numpy(), clone.sample_rate)generate_stream() yields decoded CPU float32 PCM chunks as soon as they are ready:
stream = model.generate_stream(
"The voice generated by this model sounds amazing!",
mode="voice_clone",
reference_audio="assets/default_reference.wav",
)
with sf.SoundFile(
"stream.wav",
mode="w",
samplerate=model.sample_rate,
channels=1,
subtype="PCM_16",
) as output:
for chunk in stream:
output.write(chunk.waveform.squeeze(0).numpy())cutetts --model-dir /path/to/CuteTTS --mode tts --text "The voice generated by this model sounds amazing!" --output tts.wav
cutetts --model-dir /path/to/CuteTTS-distill \
--mode voice_clone --reference-audio assets/default_reference.wav \
--text "The voice generated by this model sounds amazing!" --output clone.wav| Model | Params. | LibriSpeech test-clean WER (%) ↓ | LibriSpeech test-clean SIM ↑ | Seed-TTS EN WER (%) ↓ | Seed-TTS EN SIM ↑ | Seed-TTS ZH WER (%) ↓ | Seed-TTS ZH SIM ↑ |
|---|---|---|---|---|---|---|---|
| MOSS‑TTS | 8B | 1.98 | 67.7 | 1.84 | 70.9 | 1.37 | 77.0 |
| Qwen3‑TTS | 1.7B | 2.35 | 70.3 | 1.66 | 71.4 | 0.91 | 77.0 |
| FireRedTTS‑2 | 1.5B | 4.32 | 64.2 | 1.95 | 66.5 | 1.14 | 73.6 |
| MOSS‑TTS‑Nano | 0.1B | 4.10 | 48.4 | 4.62 | 49.9 | 3.13 | 64.3 |
| F5‑TTS | 0.3B | 2.42 | 66.0 | 1.83 | 67.0 | 1.56 | 76.0 |
| ZipVoice | 0.1B | 2.05 | 67.4 | 1.70 | 69.7 | 1.40 | 75.1 |
| IndexTTS2 | 1.5B | 2.47 | 70.0 | 2.22 | 70.6 | 1.02 | 76.5 |
| CosyVoice 3 | 0.5B | 1.99 | 69.7 | 2.02 | 71.8 | 1.16 | 78.0 |
| VoxCPM2 | 2B | 3.01 | 74.0 | 1.84 | 75.3 | 0.97 | 79.5 |
| VibeVoice | 1.5B | – | – | 3.04 | 68.9 | 1.16 | 74.4 |
| DiTAR | 0.6B | 2.39 | 67.0 | 1.69 | 73.5 | 1.02 | 75.3 |
| VibeVoice‑Realtime | 0.5B | 2.00 | 69.5 | 2.05 | 63.3 | – | – |
| Pocket TTS | 0.1B | 1.59 | 49.1 | 1.63 | 50.7 | – | – |
| CuteTTS | 0.2B | 2.16 | 78.9 | 2.04 | 76.5 | 1.41 | 77.8 |
| CuteTTS‑distill | 0.2B | 2.41 | 76.8 | 2.03 | 74.2 | 1.47 | 75.6 |
- Descript Audio Codec (DAC) for portions of the Audio VAE implementation
- F5-TTS for the Sway Sampling schedule
- Qwen3 model architecture through Hugging Face Transformers v4.51.0.
CuteTTS is released under the Apache License 2.0. Third-party portions retain their respective copyright notices and licenses as described in NOTICE.
