Orpheus TTS

An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.

FreeProprietaryCommand line

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About Orpheus TTS

Orpheus TTS is a speech synthesis model from Canopy Labs. You run it yourself from Python code, and the repository includes a PyPI package, a realtime streaming example, extra inference options and a list of supported emotion tags. Scripts for fine-tuning and pretraining are included for people who want to train their own voices.

A family of multilingual models was released in April 2025 as a research preview. It suits developers and hobbyists who want to run expressive speech generation on their own hardware rather than through a paid service.

Strengths

  • Emotion tags for expressive speech
  • Realtime streaming example included
  • Fine-tuning and pretraining code published
  • Installable as a Python package

Limitations

  • No packaged desktop app; needs a Python setup
  • Multilingual models are a research preview

Details

Pricing
FreeFree to download and run yourself.
License
Proprietary
Developer
Canopy Labs
Platforms
Command line
How it runs
Downloadable app
Best suited for
Developers who want to run expressive text-to-speech on their own hardware
Last verified
Added

Alternatives to Orpheus TTS

Compare all

Software that can replace Orpheus TTS for an important use case, and what changes if you switch.

  • Zonos

    An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.

    Zonos is an Apache-2.0 open-weight model with voice cloning from a short sample and a Gradio interface with Docker, but it is an early v0.1 release that needs a GPU.

  • Chatterbox

    Open source text-to-speech model with zero-shot voice cloning.

    Chatterbox is MIT licensed and adds zero-shot voice cloning from a very short sample, but it also lacks a GUI and is mainly used through code.

  • F5-TTS

    Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.

    F5-TTS adds voice cloning from a short reference recording and a local Gradio interface, while still needing a Python setup and ideally a capable GPU.

  • IndexTTS

    A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.

    FreeProprietaryCommand line

    IndexTTS offers zero-shot voice cloning with control over emotion and duration and a bundled web UI, but its license terms need checking before commercial use.

  • CosyVoice

    An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.

    CosyVoice is Apache-2.0 licensed with multilingual zero-shot cloning, a web UI, Docker and vLLM support, and it lists Linux and self-hosted platforms.

  • Fish Speech

    A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.

    FreeProprietarySelf-hosted

    Fish Speech is a self-hosted multilingual voice cloning model with a web UI and Docker Compose setups including AMD ROCm, and its license should be checked for commercial use.

  • Bark

    Open source generative audio model for expressive multilingual speech.

    Bark is MIT licensed and can produce nonverbal sounds like laughter, but it is slower, more resource-intensive and has had no repository update since mid-2024.

  • Kokoro

    Small, fast open-weight text-to-speech model you can run locally.

    Kokoro is a much smaller Apache-licensed model that runs fast and cheaply, trading the emotion tags and expressive focus for speed and lower hardware needs.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing