- Home
- Alternatives
- Alternatives to CosyVoice
Alternatives to CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code. The listings below can replace it for an important use case. Each note says what changes if you switch.
The original
CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.
Replacements
Listings that take over the same core job as CosyVoice.
Fish Speech
A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.
Fish Speech offers self-hosted multilingual voice cloning with a web UI and Docker including AMD ROCm, but it is not open source, so license terms need checking.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS clones voices from a short reference with a local Gradio interface on Windows, macOS and Linux, and is research code without a packaged installer.
IndexTTS
A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.
IndexTTS provides zero-shot cloning with control over emotion and speech duration and a bundled web UI, but it is not open source and needs license checks.
Zonos
An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.
Zonos is an Apache-2.0 open-weight model with voice cloning, Gradio and Docker, trained on over 200k hours of multilingual speech, though still an early v0.1 release.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI on Windows and Linux, and is not open source.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Chatterbox is an MIT-licensed model with zero-shot cloning from very short samples across Windows, macOS and Linux, but it has no dedicated GUI and is used through code.
OpenVoice
Open source instant voice cloning model with flexible style control.
OpenVoice is an MIT-licensed instant cloning model with fine control of emotion and accent, used through code, with no repository update since April 2025.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages and full training support, but has not been updated since mid-2024.
Also worth comparing
These listings name CosyVoice as their own alternative, so the relationship runs both ways.
AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
CosyVoice is an Apache-2.0 multilingual model with zero-shot cloning, training code, a web UI and vLLM support, aimed at technical users on Linux with a GPU.
MeloTTS
An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.
CosyVoice is an Apache-2.0 multilingual model that adds zero-shot voice cloning and training code, but needs a GPU for practical use.
Orpheus TTS
An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.
CosyVoice is Apache-2.0 licensed with multilingual zero-shot cloning, a web UI, Docker and vLLM support, and it lists Linux and self-hosted platforms.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
CosyVoice is an Apache licensed multilingual model with zero-shot voice cloning, training code, web UI and Docker, and is actively developed.
Similar software
Related functionality, not a direct replacement.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.
Fish Audio
A web platform for AI text-to-speech, voice cloning and speech-to-text with a large community voice library.