Zonos
An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About Zonos
Zonos-v0.1 is a text-to-speech model from Zyphra, trained on more than 200,000 hours of varied multilingual speech. It can clone a voice from a reference sample and offers conditioning controls that shape how expressive the output sounds.
The repository includes a Gradio web interface, a sample script, a Dockerfile and a docker-compose file for running it on your own machine. It suits people who want expressive speech synthesis without sending text to a hosted service.
Strengths
- Voice cloning from a short reference sample
- Trained on over 200k hours of multilingual speech
- Gradio interface and Docker setup included
- Runs locally without a hosted account
Limitations
- Needs Python setup and a capable GPU for comfortable speed
- Early v0.1 release with a small commit history
Details
- Pricing
- FreeFree open-weight model and code.
- License
- Apache-2.0
- Developer
- Zyphra
- Platforms
- Linux, Self-hosted
- How it runs
- Downloadable app, Self-hosted
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Running expressive text-to-speech and voice cloning on your own hardware
- Categories
- Local AI tools, AI voice
- Last verified
- Added
- Sources
Alternatives to Zonos
Compare allSoftware that can replace Zonos for an important use case, and what changes if you switch.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS also clones voices from a short reference with a local Gradio interface, and runs on Windows, macOS and Linux as a large, active project.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Chatterbox is MIT-licensed with zero-shot voice cloning from a very short sample across Windows, macOS and Linux, but has no dedicated GUI.
CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.
CosyVoice is Apache-2.0 licensed with multilingual zero-shot cloning, training code, a web UI, Docker and vLLM support, and an active 3.0 release.
Fish Speech
A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.
Fish Speech is self-hosted multilingual voice cloning with a web UI and Docker Compose, including AMD ROCm, though license terms need checking for commercial use.
IndexTTS
A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.
IndexTTS offers zero-shot cloning with emotion and duration control and a bundled web UI, but its license should be checked before commercial use.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI on Windows and Linux, and is not open source.
OpenVoice
Open source instant voice cloning model with flexible style control.
OpenVoice is MIT-licensed instant voice cloning with emotion and accent control across Windows, macOS and Linux, used mainly through code without an official GUI.
AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
AllTalk TTS is an AGPL-3.0 local server with voice cloning, finetuning and a JSON API that integrates with SillyTavern and other chat front ends.
Zonos as an alternative
Listings that name Zonos as an alternative.
Bark
Open source generative audio model for expressive multilingual speech.
Zonos adds voice cloning from a short sample, a Gradio interface and Docker setup under Apache-2.0, but it is an early v0.1 release for Linux.
Orpheus TTS
An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.
Zonos is an Apache-2.0 open-weight model with voice cloning from a short sample and a Gradio interface with Docker, but it is an early v0.1 release that needs a GPU.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
Zonos is an Apache licensed open-weight model with voice cloning, trained on over 200k hours of multilingual speech, with Gradio and Docker included.
Similar software
Related functionality, not necessarily a direct replacement.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.
Fish Audio
A web platform for AI text-to-speech, voice cloning and speech-to-text with a large community voice library.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.