Alternatives to Speaches

A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models. The listings below can replace it for an important use case. Each note says what changes if you switch.

The original

Replacements

Listings that take over the same core job as Speaches.

  • Kokoro-FastAPI

    A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.

    Kokoro-FastAPI provides an OpenAI-compatible speech endpoint in prebuilt Docker images, but it covers only Kokoro text-to-speech, not transcription or translation.

  • AllTalk TTS

    A local text-to-speech server with voice cloning, model finetuning, a settings page and a JSON API.

    AllTalk TTS is an AGPL-3.0 local text-to-speech server with a JSON API, voice cloning and finetuning, but it does not provide speech-to-text.

  • Cartesia

    An AI voice platform with an API for real-time text-to-speech, speech-to-text and voice agents.

    PaidProprietaryWeb

    Cartesia is a paid cloud API covering text-to-speech, transcription and voice agents with low latency, replacing local models with a hosted service.

  • ElevenLabs

    An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.

    FreemiumProprietaryWeb

    ElevenLabs offers text to speech, speech to text and voice cloning through an API and SDKs, but it runs only as an online service.

Similar software

Related functionality, not a direct replacement.