MeloTTS

An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About MeloTTS

MeloTTS is a text-to-speech library from MIT and MyShell.ai that supports English in American, British, Indian and Australian accents, as well as Spanish, French, Chinese mixed with English, Japanese and Korean.

It can be used from Python, and the repository also includes a Dockerfile and documentation for running it locally. It suits developers and hobbyists who want lightweight, self-hosted speech synthesis in several languages.

Strengths

  • Several English accents plus five other languages
  • Runs locally with a Dockerfile provided
  • Chinese voice handles mixed English text

Limitations

  • No voice cloning
  • Requires Python or Docker setup
  • Development has slowed

Details

Pricing
FreeFree and open source.
License
MIT
Developer
MyShell.ai
Platforms
Linux, Self-hosted, Command line
How it runs
Downloadable app, Self-hosted
Account
Not required
Works offline
Yes
Best suited for
Lightweight local text-to-speech in several languages
Last verified
Added

Alternatives to MeloTTS

Compare all

Software that can replace MeloTTS for an important use case, and what changes if you switch.

  • Piper

    A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.

    Piper is a GPL-3.0 local engine fast enough for a Raspberry Pi, running on Windows, macOS and Linux through the command line, Python or a C library.

  • Kokoro

    Small, fast open-weight text-to-speech model you can run locally.

    Kokoro is a small, fast Apache-2.0 open-weight model run locally as a Python library on Windows, macOS and Linux, without MeloTTS's web UI.

  • Kokoro-FastAPI

    A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.

    Kokoro-FastAPI serves the Kokoro model through a Docker server with an OpenAI-compatible endpoint, voice mixing, SSML and caption timestamps.

  • Coqui TTS

    Deep learning toolkit for training and running text-to-speech models.

    Coqui TTS is an MPL-2.0 toolkit with pretrained models in many more languages and training support, but it has not been updated since mid-2024.

  • RHVoice

    A free, open-source speech synthesizer for Russian and other languages, used with screen readers and desktops.

    RHVoice is a lightweight open-source synthesizer for Linux and Android with strong Russian support, sounding less natural than neural text-to-speech.

  • CosyVoice

    An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.

    CosyVoice is an Apache-2.0 multilingual model that adds zero-shot voice cloning and training code, but needs a GPU for practical use.

  • Tortoise TTS

    Open source multi-voice text-to-speech system tuned for quality.

    Tortoise TTS is an Apache-2.0 multi-voice system focused on realistic prosody, notably slower to generate than MeloTTS and not updated since late 2024.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing