Alternatives to Bark

Open source generative audio model for expressive multilingual speech. The listings below can replace it for an important use case. Each note says what changes if you switch.

The original

Replacements

Listings that take over the same core job as Bark.

  • Chatterbox

    Open source text-to-speech model with zero-shot voice cloning.

    Chatterbox is an MIT licensed text-to-speech model focused on zero-shot voice cloning from a short sample, used mainly through code without a dedicated GUI.

  • Kokoro

    Small, fast open-weight text-to-speech model you can run locally.

    Kokoro is a small Apache-licensed open-weight model that runs fast and cheaply, used as a Python library, trading Bark's nonverbal sounds for speed.

  • Coqui TTS

    Deep learning toolkit for training and running text-to-speech models.

    Coqui TTS is an MPL-2.0 toolkit with pretrained models in many languages plus training and fine-tuning, though updates stopped after the company shut down.

  • Tortoise TTS

    Open source multi-voice text-to-speech system tuned for quality.

    Tortoise TTS is an Apache-2.0 multi-voice system that emphasizes realistic prosody and intonation but is notably slower to generate speech.

  • OpenVoice

    Open source instant voice cloning model with flexible style control.

    OpenVoice is an MIT licensed voice cloning model with fine control over emotion and accent, used through code, with no updates since April 2025.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS adds voice cloning from a short reference recording and a local Gradio interface, but needs a Python setup and ideally a capable GPU.

  • Piper

    A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.

    Piper is a fast GPL-3.0 TTS engine that runs on low-power hardware like a Raspberry Pi, trading expressive generation for speed and command-line integration.

  • GPT-SoVITS

    A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.

    FreeProprietaryWindowsLinux

    GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI, runs on Windows and Linux, and is not open source.

Similar software

Related functionality, not a direct replacement.