Kokoro-FastAPI
A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About Kokoro-FastAPI
Kokoro-FastAPI wraps the Kokoro-82M text-to-speech model in an OpenAI-compatible speech endpoint. It supports English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin Chinese, with voice mixing, SSML, and multi-speaker generation.
It can produce per-word or per-chunk caption timestamps and has phoneme endpoints. Prebuilt images cover CPU and NVIDIA GPUs, with experimental AMD ROCm. Apple Silicon works when run directly with UV. An optional web UI offers read-along playback of long generations.
Strengths
- OpenAI-compatible speech endpoint
- Prebuilt multi-architecture Docker images with models included
- Caption timestamps and phoneme endpoints
- Voice mixing and SSML support
Limitations
- AMD GPU support is experimental
- Apple Silicon has no prebuilt image
Details
- Pricing
- FreeFree and open source under the Apache 2.0 licence.
- License
- Apache-2.0
- Developer
- remsky
- Platforms
- macOS, Linux, Self-hosted
- How it runs
- Self-hosted
- Account
- Not required
- Best suited for
- Self-hosting a fast text-to-speech back end for apps, readers or Home Assistant
- Categories
- Local AI tools, AI voice, AI developer tools
- Last verified
- Added
Alternatives to Kokoro-FastAPI
Compare allSoftware that can replace Kokoro-FastAPI for an important use case, and what changes if you switch.
Speaches
A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.
Speaches is an MIT-licensed self-hosted OpenAI-compatible server that adds speech-to-text and streaming transcription, and serves both Kokoro and Piper voices through Docker.
Kokoro
Small, fast open-weight text-to-speech model you can run locally.
Kokoro is the underlying Apache-2.0 model used directly as a Python library, so moving drops the Docker server, OpenAI-compatible endpoint and caption timestamps.
AllTalk TTS
A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.
AllTalk TTS is an AGPL-3.0 local server with a JSON API, voice cloning and finetuning on Windows and Linux, rather than an OpenAI-compatible endpoint.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Piper is a GPL-3.0 local engine fast enough for a Raspberry Pi, used from the command line, Python or a C library instead of an HTTP server.
MeloTTS
An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.
MeloTTS is an MIT-licensed multilingual engine with a command line, web UI and Dockerfile, offering several English accents and five other languages, though development has slowed.
Cartesia
An AI voice platform with an API for real-time text-to-speech, speech-to-text and voice agents.
Cartesia is a paid cloud API for real-time text-to-speech, transcription and voice agents, replacing your own hardware with a hosted service and no local option.
Similar software
Related functionality, not necessarily a direct replacement.
Speech Note
A Linux app for offline speech-to-text, text-to-speech and machine translation using local models.
Whisper
An open-source speech recognition model from OpenAI for transcribing and translating audio on your own machine.
whisper.cpp
C/C++ port of OpenAI's Whisper for fast offline speech-to-text on your own hardware.
Chatterbox
Open source text-to-speech model with zero-shot voice cloning.
Orpheus TTS
An open text-to-speech model built on a language model, aimed at natural, emotionally expressive speech.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.