IndexTTS
A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About IndexTTS
IndexTTS generates speech in a new voice from a short reference sample, without training on that speaker first. The project describes it as an industrial-level, controllable and efficient system, with options to steer the emotion and the duration of the generated speech.
The repository includes the Python package, a web UI script for trying it in a browser on your own machine, and documentation in several languages. It suits developers and hobbyists who want to run voice cloning and speech synthesis locally rather than through a cloud API.
Strengths
- Zero-shot voice cloning from a reference sample
- Control over emotion and duration
- Runs locally with a bundled web UI
- Large and active project on GitHub
Limitations
- Requires a Python setup and suitable hardware
- License terms need checking before commercial use
Details
- Pricing
- FreeFree to download and run from the GitHub repository.
- License
- Proprietary
- Developer
- The IndexTTS contributors
- Platforms
- Command line
- How it runs
- Downloadable app
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Running voice cloning and text-to-speech on your own machine
- Categories
- Local AI tools, AI voice
- Last verified
- Added
Alternatives to IndexTTS
Compare allSoftware that can replace IndexTTS for an important use case, and what changes if you switch.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS clones voices from a short reference with a local Gradio interface on Windows, macOS and Linux, without IndexTTS's explicit duration control.
CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.
CosyVoice is an Apache-2.0 multilingual model with zero-shot cloning, training code, a web UI and vLLM support, avoiding the license checks IndexTTS needs.
Fish Speech
A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.
Fish Speech is a self-hosted multilingual cloning model with a web UI and Docker Compose setups including AMD ROCm, and its license also needs checking.
Zonos
An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.
Zonos is an Apache-2.0 open-weight model with cloning from a short sample, a Gradio interface and Docker on Linux, in an early v0.1 release.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI and API scripts on Windows and Linux, using few-shot training rather than zero-shot.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Chatterbox is an MIT-licensed zero-shot cloning model for Windows, macOS and Linux, with clearer licensing but no dedicated GUI.
OpenVoice
Open source instant voice cloning model with flexible style control.
OpenVoice is an MIT-licensed instant cloning model with control of emotion and accent, used through code, with no repository update since April 2025.
Orpheus TTS
An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.
Orpheus TTS uses emotion tags for expressive speech and includes streaming and fine-tuning code, but its multilingual models are a research preview.
Similar software
Related functionality, not necessarily a direct replacement.
AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.
Bark
Open source generative audio model for expressive multilingual speech.