Open-source text-to-speech toolkit with XTTS voice cloning that runs anywhere.
Coqui is the open-source speech toolkit behind XTTS, one of the best free voice-cloning models available. With a few seconds of reference audio you can synthesize natural speech in 17+ languages, entirely on your own hardware. It is the go-to for indie developers, accessibility projects, and anyone who does not want to pay per character for TTS.
Who it's for: Developers and makers who need multilingual, cloneable speech without per-use fees or cloud lock-in.
Clone a voice from a short sample in many languages.
17+ languages with a single model.
Run inference on your own GPU or CPU.
Permissively licensed for commercial use.
If you can run Python, Coqui gives you near-commercial TTS for free. It is the backbone of countless open-source voice projects.