Alternatives to Bark
Open source generative audio model for expressive multilingual speech. The listings below can replace it for an important use case. Each note says what changes if you switch.
The original
Bark
Open source generative audio model for expressive multilingual speech.
Replacements
Listings that take over the same core job as Bark.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Chatterbox is an MIT licensed text-to-speech model focused on zero-shot voice cloning from a short sample, used mainly through code without a dedicated GUI.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
Kokoro is a small Apache-licensed open-weight model that runs fast and cheaply, used as a Python library, trading Bark's nonverbal sounds for speed.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages plus training and fine-tuning, though updates stopped after the company shut down.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
Tortoise TTS is an Apache-2.0 multi-voice system that emphasizes realistic prosody and intonation but is notably slower to generate speech.
OpenVoice
Open source instant voice cloning model with flexible style control.
OpenVoice is an MIT licensed voice cloning model with fine control over emotion and accent, used through code, with no updates since April 2025.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS adds voice cloning from a short reference recording and a local Gradio interface, but needs a Python setup and ideally a capable GPU.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Piper is a fast GPL-3.0 TTS engine that runs on low-power hardware like a Raspberry Pi, trading expressive generation for speed and command-line integration.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI, runs on Windows and Linux, and is not open source.
Similar software
Related functionality, not a direct replacement.
ACE-Step
Open source local music generation model that runs on your own hardware.
Whisper
An open-source speech recognition model from OpenAI for transcribing and translating audio on your own machine.
Applio
A free, open-source AI voice conversion suite for covers, custom voice training and real-time voice changing.