CosyVoice

An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About CosyVoice

CosyVoice is a multilingual voice generation model with code for inference, training and deployment. Three generations are published: CosyVoice 1.0, 2.0 and Fun-CosyVoice 3.0, with weights on ModelScope and Hugging Face.

The repository includes a web UI script, Docker files, a runtime for deployment and a vLLM example. It suits developers and enthusiasts who want to run text-to-speech and voice cloning locally.

Strengths

  • Zero-shot voice cloning across multiple languages
  • Includes training and deployment code, not only inference
  • Web UI, Docker and vLLM support
  • Actively developed with a 3.0 release

Limitations

  • Requires Python setup and a GPU for practical use
  • Aimed at technical users rather than casual listeners

Details

Pricing
FreeFree and open source.
License
Apache-2.0
Developer
Alibaba
Platforms
Linux, Self-hosted
How it runs
Downloadable app, Self-hosted
Account
Not required
Works offline
Yes
Best suited for
Self-hosted multilingual text-to-speech and voice cloning
Last verified
Added

Alternatives to CosyVoice

Compare all

Software that can replace CosyVoice for an important use case, and what changes if you switch.

  • Fish Speech

    A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.

    FreeProprietarySelf-hosted

    Fish Speech offers self-hosted multilingual voice cloning with a web UI and Docker including AMD ROCm, but it is not open source, so license terms need checking.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS clones voices from a short reference with a local Gradio interface on Windows, macOS and Linux, and is research code without a packaged installer.

  • IndexTTS

    A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.

    FreeProprietaryCommand line

    IndexTTS provides zero-shot cloning with control over emotion and speech duration and a bundled web UI, but it is not open source and needs license checks.

  • Zonos

    An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.

    Zonos is an Apache-2.0 open-weight model with voice cloning, Gradio and Docker, trained on over 200k hours of multilingual speech, though still an early v0.1 release.

  • GPT-SoVITS

    A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.

    FreeProprietaryWindowsLinux

    GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI on Windows and Linux, and is not open source.

  • Chatterbox

    Open source text-to-speech model with zero-shot voice cloning.

    Chatterbox is an MIT-licensed model with zero-shot cloning from very short samples across Windows, macOS and Linux, but it has no dedicated GUI and is used through code.

  • OpenVoice

    Open source instant voice cloning model with flexible style control.

    OpenVoice is an MIT-licensed instant cloning model with fine control of emotion and accent, used through code, with no repository update since April 2025.

  • Coqui TTS

    Deep learning toolkit for training and running text-to-speech models.

    Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages and full training support, but has not been updated since mid-2024.

CosyVoice as an alternative

Listings that name CosyVoice as an alternative.

  • AllTalk TTS

    A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.

    CosyVoice is an Apache-2.0 multilingual model with zero-shot cloning, training code, a web UI and vLLM support, aimed at technical users on Linux with a GPU.

  • MeloTTS

    An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.

    CosyVoice is an Apache-2.0 multilingual model that adds zero-shot voice cloning and training code, but needs a GPU for practical use.

  • Orpheus TTS

    An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.

    FreeProprietaryCommand line

    CosyVoice is Apache-2.0 licensed with multilingual zero-shot cloning, a web UI, Docker and vLLM support, and it lists Linux and self-hosted platforms.

  • Tortoise TTS

    Open source multi-voice text-to-speech system tuned for quality.

    CosyVoice is an Apache licensed multilingual model with zero-shot voice cloning, training code, web UI and Docker, and is actively developed.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing