AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About AllTalk TTS
AllTalk TTS is a text-to-speech server based on the Coqui TTS engine. It runs standalone or as a voice back end for Text-generation-webui, SillyTavern and KoboldCPP, and other programs can call it through a JSON API.
It can finetune XTTSv2 models on a chosen voice, generate long audio in bulk, and use separate voices for characters and narration. DeepSpeed support and a low VRAM mode help on smaller GPUs. Version 2 is the recommended release and is where current development happens.
Strengths
- Voice finetuning on your own samples
- Integrates with SillyTavern, KoboldCPP and Text-generation-webui
- Low VRAM mode and DeepSpeed support
- Bulk generator for long narration
Limitations
- Version 1 and version 2 are documented separately, which can confuse new users
- Setup utility covers Windows and Linux only
Details
- Pricing
- FreeFree and open source under the AGPL-3.0 licence.
- License
- AGPL-3.0
- Developer
- erew123
- Platforms
- Windows, Linux, Self-hosted
- How it runs
- Downloadable app, Self-hosted
- Account
- Not required
- Best suited for
- Local voice output and voice cloning for AI chat front ends
- Categories
- Local AI tools, AI voice
- Last verified
- Added
Alternatives to AllTalk TTS
Compare allSoftware that can replace AllTalk TTS for an important use case, and what changes if you switch.
Kokoro-FastAPI
A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.
Kokoro-FastAPI is an Apache-2.0 Docker server with an OpenAI-compatible speech endpoint and voice mixing, but it serves the Kokoro model and does not offer voice cloning or finetuning.
Speaches
A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.
Speaches is an MIT-licensed self-hosted server with an OpenAI-style API that covers both speech-to-text and text-to-speech with Kokoro and Piper voices, set up through Docker.
Coqui TTS
Deep learning toolkit for training and running text-to-speech models.
Coqui TTS is an MPL-2.0 toolkit for running and training speech models in many languages, used from code rather than a settings page, and has not been updated since mid-2024.
Fish Speech
A multilingual text-to-speech and voice cloning model you run yourself, with Docker images and a web UI.
Fish Speech offers self-hosted multilingual voice cloning with a web UI and Docker setups including AMD ROCm, but it is not open source and its license needs checking for commercial use.
GPT-SoVITS
A voice cloning and text-to-speech toolkit that can train a voice from about one minute of audio.
GPT-SoVITS trains a custom voice from about a minute of audio with a local web UI and API scripts on Windows and Linux, though it is not open source.
CosyVoice
An open-source multilingual text-to-speech model with zero-shot voice cloning, inference and training code.
CosyVoice is an Apache-2.0 multilingual model with zero-shot cloning, training code, a web UI and vLLM support, aimed at technical users on Linux with a GPU.
F5-TTS
Open-source text-to-speech model that can clone voices, with a local Gradio interface for running it.
F5-TTS clones voices from a short reference recording through a local Gradio interface on Windows, macOS and Linux, but it lacks AllTalk's integrations with chat front ends.
Zonos
An open-weight text-to-speech model with voice cloning that you run locally through a Gradio interface.
Zonos is an Apache-2.0 open-weight model with voice cloning, a Gradio interface and Docker setup for Linux, released as an early v0.1 with a small commit history.
AllTalk TTS as an alternative
Listings that name AllTalk TTS as an alternative.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.
AllTalk TTS is a free AGPL-3.0 local server with voice cloning and finetuning on Windows and Linux, replacing a hosted service with your own hardware.
Similar software
Related functionality, not necessarily a direct replacement.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Tortoise TTS
Open source multi-voice text-to-speech system tuned for quality.
IndexTTS
A zero-shot text-to-speech system you run yourself, with control over emotion and speech duration.
Applio
A free, open-source AI voice conversion suite for covers, custom voice training and real-time voice changing.