Alternatives to Chatterbox
Open source text-to-speech model with zero-shot voice cloning. The listings below can replace it for an important use case. Each note says what changes if you switch.
The original
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Replacements
Listings that take over the same core job as Chatterbox.
OpenVoice
Open source instant voice cloning model with flexible style control.
OpenVoice is an MIT licensed voice cloning model that adds fine control over emotion and accent, used through code, with no updates since April 2025.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages and full training support, though its repository stopped updating in mid-2024.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
Tortoise TTS is an Apache-2.0 multi-voice system focused on realistic prosody, notably slower than newer models and not updated since late 2024.
Bark
Open source generative audio model for expressive multilingual speech.
Bark is an MIT licensed generative audio model that adds nonverbal sounds like laughter, but it is slower and more resource-intensive than lightweight options.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
Kokoro is a small Apache-licensed open-weight text-to-speech model that runs fast, used as a Python library, without zero-shot voice cloning listed.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS also clones voices from a short reference recording and adds a local Gradio interface, but needs a Python setup and ideally a capable GPU.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a voice from about a minute of audio with a local web UI and API scripts, runs on Windows and Linux, and is not open source.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Piper is a fast GPL-3.0 TTS engine that runs even on a Raspberry Pi, suited to offline integration, but relies on prebuilt voice models rather than cloning.
Similar software
Related functionality, not a direct replacement.
Whisper
An open-source speech recognition model from OpenAI for transcribing and translating audio on your own machine.
ACE-Step
Open source local music generation model that runs on your own hardware.
Applio
A free, open-source AI voice conversion suite for covers, custom voice training and real-time voice changing.