Alternatives to Fish Speech

A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI. The listings below can replace it for an important use case. Each note says what changes if you switch.

The original

  • Fish Speech

    A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.

    FreeProprietarySelf-hosted

Replacements

Listings that take over the same core job as Fish Speech.

  • CosyVoice

    An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.

    CosyVoice is Apache-2.0 licensed with zero-shot multilingual cloning, training code, a web UI and vLLM support, avoiding Fish Speech's license questions but needing Linux and a GPU.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS clones voices from a short reference with a local Gradio interface on Windows, macOS and Linux, and is research code without a packaged installer.

  • IndexTTS

    A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.

    FreeProprietaryCommand line

    IndexTTS offers zero-shot cloning with control over emotion and duration and a bundled web UI, with license terms that also need checking for commercial use.

  • Zonos

    An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.

    Zonos is an Apache-2.0 open-weight model with cloning, a Gradio interface and Docker setup, though it is an early v0.1 release with a small commit history.

  • GPT-SoVITS

    A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.

    FreeProprietaryWindowsLinux

    GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI, API scripts and Docker or Colab options on Windows and Linux.

  • Chatterbox

    Open source text-to-speech model with zero-shot voice cloning.

    Chatterbox is an MIT-licensed model with zero-shot cloning from very short samples, usable locally on Windows, macOS and Linux but without a dedicated GUI.

  • AllTalk TTS

    A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.

    AllTalk TTS is an AGPL-3.0 local server with cloning, finetuning, a settings page and JSON API, integrated with SillyTavern and other chat front ends.

  • Fish Audio

    A web platform for AI text-to-speech, voice cloning and speech-to-text with a large community voice library.

    FreemiumProprietaryWeb

    Fish Audio is a freemium hosted platform for text-to-speech and voice cloning with a large community voice library, removing the GPU setup but running only online.

Also worth comparing

These listings name Fish Speech as their own alternative, so the relationship runs both ways.

  • ElevenLabs

    An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.

    FreemiumProprietaryWeb

    Fish Speech runs multilingual voice cloning on your own GPU with a web UI and Docker, avoiding the cloud but needing setup and a license check.

  • OpenVoice

    Open source instant voice cloning model with flexible style control.

    Fish Speech is a self-hosted multilingual cloning model with a web UI and Docker, needing a capable GPU and a license check before commercial use.

  • Orpheus TTS

    An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.

    FreeProprietaryCommand line

    Fish Speech is a self-hosted multilingual voice cloning model with a web UI and Docker Compose setups including AMD ROCm, and its license should be checked for commercial use.

Similar software

Related functionality, not a direct replacement.