Alternatives to AllTalk TTS

A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API. The listings below can replace it for an important use case. Each note says what changes if you switch.

The original

Replacements

Listings that take over the same core job as AllTalk TTS.

  • Kokoro-FastAPI

    A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.

    Kokoro-FastAPI is an Apache-2.0 Docker server with an OpenAI-compatible speech endpoint and voice mixing, but it serves the Kokoro model and does not offer voice cloning or finetuning.

  • Speaches

    A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.

    Speaches is an MIT-licensed self-hosted server with an OpenAI-style API that covers both speech-to-text and text-to-speech with Kokoro and Piper voices, set up through Docker.

  • Coqui TTS

    Deep learning toolkit for training and running text-to-speech models.

    Coqui TTS is an MPL-2.0 toolkit for running and training speech models in many languages, used from code rather than a settings page, and has not been updated since mid-2024.

  • Fish Speech

    A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.

    FreeProprietarySelf-hosted

    Fish Speech offers self-hosted multilingual voice cloning with a web UI and Docker setups including AMD ROCm, but it is not open source and its license needs checking for commercial use.

  • GPT-SoVITS

    A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.

    FreeProprietaryWindowsLinux

    GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI and API scripts on Windows and Linux, though it is not open source.

  • CosyVoice

    An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.

    CosyVoice is an Apache-2.0 multilingual model with zero-shot cloning, training code, a web UI and vLLM support, aimed at technical users on Linux with a GPU.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS clones voices from a short reference recording through a local Gradio interface on Windows, macOS and Linux, but it lacks AllTalk's integrations with chat front ends.

  • Zonos

    An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.

    Zonos is an Apache-2.0 open-weight model with voice cloning, a Gradio interface and Docker setup for Linux, released as an early v0.1 with a small commit history.

Also worth comparing

These listings name AllTalk TTS as their own alternative, so the relationship runs both ways.

  • ElevenLabs

    An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.

    FreemiumProprietaryWeb

    AllTalk TTS is a free AGPL-3.0 local server with voice cloning and finetuning on Windows and Linux, replacing a hosted service with your own hardware.

Similar software

Related functionality, not a direct replacement.