Kokoro-FastAPI logo

Kokoro-FastAPI

A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About Kokoro-FastAPI

Kokoro-FastAPI wraps the Kokoro-82M text-to-speech model in an OpenAI-compatible speech endpoint. It supports English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin Chinese, with voice mixing, SSML, and multi-speaker generation.

It can produce per-word or per-chunk caption timestamps and has phoneme endpoints. Prebuilt images cover CPU and NVIDIA GPUs, with experimental AMD ROCm. Apple Silicon works when run directly with UV. An optional web UI offers read-along playback of long generations.

Strengths

  • OpenAI-compatible speech endpoint
  • Prebuilt multi-architecture Docker images with models included
  • Caption timestamps and phoneme endpoints
  • Voice mixing and SSML support

Limitations

  • AMD GPU support is experimental
  • Apple Silicon has no prebuilt image

Details

Pricing
FreeFree and open source under the Apache 2.0 licence.
License
Apache-2.0
Developer
remsky
Platforms
macOS, Linux, Self-hosted
How it runs
Self-hosted
Account
Not required
Best suited for
Self-hosting a fast text-to-speech back end for apps, readers or Home Assistant
Last verified
Added

Alternatives to Kokoro-FastAPI

Compare all

Software that can replace Kokoro-FastAPI for an important use case, and what changes if you switch.

  • Speaches

    A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.

    Speaches is an MIT-licensed self-hosted OpenAI-compatible server that adds speech-to-text and streaming transcription, and serves both Kokoro and Piper voices through Docker.

  • Kokoro

    Small, fast open-weight text-to-speech model you can run locally.

    Kokoro is the underlying Apache-2.0 model used directly as a Python library, so moving drops the Docker server, OpenAI-compatible endpoint and caption timestamps.

  • AllTalk TTS

    A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.

    AllTalk TTS is an AGPL-3.0 local server with a JSON API, voice cloning and finetuning on Windows and Linux, rather than an OpenAI-compatible endpoint.

  • Piper

    A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.

    Piper is a GPL-3.0 local engine fast enough for a Raspberry Pi, used from the command line, Python or a C library instead of an HTTP server.

  • MeloTTS

    An open-source multilingual text-to-speech engine from MIT and MyShell.ai with a command line and web UI.

    MeloTTS is an MIT-licensed multilingual engine with a command line, web UI and Dockerfile, offering several English accents and five other languages, though development has slowed.

  • Cartesia

    An AI voice platform with an API for real-time text-to-speech, speech-to-text and voice agents.

    PaidProprietaryWeb

    Cartesia is a paid cloud API for real-time text-to-speech, transcription and voice agents, replacing your own hardware with a hosted service and no local option.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing