Alternatives to Chatterbox

Open source text-to-speech model with zero-shot voice cloning. The listings below can replace it for an important use case. Each note says what changes if you switch.

The original

Replacements

Listings that take over the same core job as Chatterbox.

  • OpenVoice

    Open source instant voice cloning model with flexible style control.

    OpenVoice is an MIT licensed voice cloning model that adds fine control over emotion and accent, used through code, with no updates since April 2025.

  • Coqui TTS

    Deep learning toolkit for training and running text-to-speech models.

    Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages and full training support, though its repository stopped updating in mid-2024.

  • Tortoise TTS

    Open source multi-voice text-to-speech system tuned for quality.

    Tortoise TTS is an Apache-2.0 multi-voice system focused on realistic prosody, notably slower than newer models and not updated since late 2024.

  • Bark

    Open source generative audio model for expressive multilingual speech.

    Bark is an MIT licensed generative audio model that adds nonverbal sounds like laughter, but it is slower and more resource-intensive than lightweight options.

  • Kokoro

    Small, fast open-weight text-to-speech model you can run locally.

    Kokoro is a small Apache-licensed open-weight text-to-speech model that runs fast, used as a Python library, without zero-shot voice cloning listed.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS also clones voices from a short reference recording and adds a local Gradio interface, but needs a Python setup and ideally a capable GPU.

  • GPT-SoVITS

    A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.

    FreeProprietaryWindowsLinux

    GPT-SoVITS trains a voice from about a minute of audio with a local web UI and API scripts, runs on Windows and Linux, and is not open source.

  • Piper

    A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.

    Piper is a fast GPL-3.0 TTS engine that runs even on a Raspberry Pi, suited to offline integration, but relies on prebuilt voice models rather than cloning.

Similar software

Related functionality, not a direct replacement.