Alternatives to Google Cloud Text-to-Speech
A Google Cloud API that converts text into synthetic speech with hundreds of voices and languages. Notes say what changes if you switch.
The original
Google Cloud Text-to-Speech
A Google Cloud API that converts text into synthetic speech with hundreds of voices and languages.
14 alternatives
Most similar to Google Cloud Text-to-Speech first.
Azure AI Speech Studio
Microsoft's web studio for creating text-to-speech audio and custom neural voices with Azure AI Speech.
Azure AI Speech Studio moves you to Microsoft Azure with a no-code studio and custom neural voice creation, requiring an Azure subscription.
Amazon Polly
AWS cloud text-to-speech service with neural and generative voices in dozens of languages.
Amazon Polly moves you to AWS with neural and generative voices across dozens of languages and SSML control, requiring an AWS account and billing.
Cartesia
An AI voice platform with an API for real-time text-to-speech, speech-to-text and voice agents.
Cartesia is a paid cloud API built for low-latency real-time speech, combining text-to-speech, transcription and voice agents in one service.
Deepgram
Speech-to-text, text-to-speech and voice agent APIs with a web playground for trying the models.
Deepgram pairs text-to-speech with speech-to-text and voice agent APIs, adding a free tier, browser playground and a self-hosted option.
AssemblyAI
A speech AI platform with APIs for transcription, speech understanding, dictation and voice agents.
Speechmatics
Speech recognition and text-to-speech APIs for multilingual, multi-speaker transcription and voice agents.
Speechmatics offers text-to-speech next to multilingual transcription APIs, with cloud, on-premises and on-device deployment options.
ElevenLabs
An AI voice platform for text to speech, voice cloning, dubbing, speech to text and music.
ElevenLabs is a freemium online service with voice cloning and dubbing in over 70 languages, plus an API and SDKs for developers.
edge-tts
A command-line tool and Python module that creates speech audio using Microsoft Edge's online voices.
Vapi
A developer platform for building, testing and deploying AI voice agents with telephony and integrations.
Dictation
A free web app that turns speech into text in Google Chrome, with voice commands for punctuation.
Speaches
A self-hosted, OpenAI API-compatible server for local speech-to-text, translation and text-to-speech models.
Kokoro-FastAPI
A Dockerized, OpenAI-compatible server for running the Kokoro-82M text-to-speech model on your own hardware.
Kokoro-FastAPI is a free self-hosted Docker server exposing an OpenAI-compatible speech endpoint, running the Kokoro model on your own hardware.
Sherpa-ONNX
An offline speech toolkit for speech-to-text, text-to-speech, diarization and voice activity detection on many devices.
Piper
A fast neural text-to-speech engine that runs locally, even on modest hardware like a Raspberry Pi.
Piper is a free GPL-3.0 engine running locally even on a Raspberry Pi, removing cloud billing but with fewer voices and no managed API.