Alternatives to Google Cloud Text-to-Speech

A Google Cloud API that converts text into synthetic speech with hundreds of voices and languages. Notes say what changes if you switch.

The original

14 alternatives

Most similar to Google Cloud Text-to-Speech first.

  • Azure AI Speech Studio

    Microsoft's web studio for creating text-to-speech audio and custom neural voices with Azure AI Speech.

    PaidProprietaryWeb

    Azure AI Speech Studio moves you to Microsoft Azure with a no-code studio and custom neural voice creation, requiring an Azure subscription.

  • Amazon Polly

    AWS cloud text-to-speech service with neural and generative voices in dozens of languages.

    PaidProprietaryWeb

    Amazon Polly moves you to AWS with neural and generative voices across dozens of languages and SSML control, requiring an AWS account and billing.

  • Cartesia

    An AI voice platform with an API for real-time text-to-speech, speech-to-text and voice agents.

    PaidProprietaryWeb

    Cartesia is a paid cloud API built for low-latency real-time speech, combining text-to-speech, transcription and voice agents in one service.

  • Deepgram

    Speech-to-text, text-to-speech and voice agent APIs with a web playground for trying the models.

    Deepgram pairs text-to-speech with speech-to-text and voice agent APIs, adding a free tier, browser playground and a self-hosted option.

  • AssemblyAI

    A speech AI platform with APIs for transcription, speech understanding, dictation and voice agents.

    PaidProprietaryWeb
  • Speechmatics

    Speech recognition and text-to-speech APIs for multilingual, multi-speaker transcription and voice agents.

    FreemiumProprietaryWeb

    Speechmatics offers text-to-speech next to multilingual transcription APIs, with cloud, on-premises and on-device deployment options.

  • ElevenLabs

    An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.

    FreemiumProprietaryWeb

    ElevenLabs is a freemium online service with voice cloning and dubbing in over 70 languages, plus an API and SDKs for developers.

  • edge-tts

    A command-line tool and Python module that creates speech audio using Microsoft Edge's online voices.

  • Vapi

    A developer platform for building, testing and deploying AI voice agents with telephony and integrations.

    PaidProprietaryWeb
  • Dictation

    A free web app that turns speech into text in Google Chrome, with voice commands for punctuation.

    FreeProprietaryWeb
  • Speaches

    A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.

  • Kokoro-FastAPI

    A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.

    Kokoro-FastAPI is a free self-hosted Docker server exposing an OpenAI-compatible speech endpoint, running the Kokoro model on your own hardware.

  • Sherpa-ONNX

    An offline speech toolkit for speech-to-text, text-to-speech, diarization and voice activity detection on many devices.

  • Piper

    A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.

    Piper is a free GPL-3.0 engine running locally even on a Raspberry Pi, removing cloud billing but with fewer voices and no managed API.