MeloTTS
An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About MeloTTS
MeloTTS is a text-to-speech library from MIT and MyShell.ai that supports English in American, British, Indian and Australian accents, as well as Spanish, French, Chinese mixed with English, Japanese and Korean.
It can be used from Python, and the repository also includes a Dockerfile and documentation for running it locally. It suits developers and hobbyists who want lightweight, self-hosted speech synthesis in several languages.
Strengths
- Several English accents plus five other languages
- Runs locally with a Dockerfile provided
- Chinese voice handles mixed English text
Limitations
- No voice cloning
- Requires Python or Docker setup
- Development has slowed
Details
- Pricing
- FreeFree and open source.
- License
- MIT
- Developer
- MyShell.ai
- Platforms
- Linux, Self-hosted, Command line
- How it runs
- Downloadable app, Self-hosted
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Lightweight local text-to-speech in several languages
- Categories
- Local AI tools, AI voice
- Last verified
- Added
Alternatives to MeloTTS
Compare allSoftware that can replace MeloTTS for an important use case, and what changes if you switch.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Piper is a GPL-3.0 local engine fast enough for a Raspberry Pi, running on Windows, macOS and Linux through the command line, Python or a C library.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
Kokoro is a small, fast Apache-2.0 open-weight model run locally as a Python library on Windows, macOS and Linux, without MeloTTS's web UI.
Kokoro-FastAPI
A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.
Kokoro-FastAPI serves the Kokoro model through a Docker server with an OpenAI-compatible endpoint, voice mixing, SSML and caption timestamps.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Coqui TTS is an MPL-2.0 toolkit with pretrained models in many more languages and training support, but it has not been updated since mid-2024.
RHVoice
A free, open-source speech synthesizer for Russian and other languages, used with screen readers and desktops.
RHVoice is a lightweight open-source synthesizer for Linux and Android with strong Russian support, sounding less natural than neural text-to-speech.
CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.
CosyVoice is an Apache-2.0 multilingual model that adds zero-shot voice cloning and training code, but needs a GPU for practical use.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
Tortoise TTS is an Apache-2.0 multi-voice system focused on realistic prosody, notably slower to generate than MeloTTS and not updated since late 2024.
Similar software
Related functionality, not necessarily a direct replacement.
Speaches
A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.
AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
Speech Note
A Linux app for offline speech-to-text, text-to-speech and machine translation using local models.
Balabolka
A free Windows program that reads text aloud with installed voices and saves it as audio files.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
Orpheus TTS
An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.