WhisperX

A speech recognition tool that adds word-level timestamps and speaker diarization to Whisper transcription.

FreeProprietaryCommand line

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About WhisperX

WhisperX transcribes audio using Whisper models and aligns the output to produce accurate timestamps for each word. It can also perform speaker diarization, labelling which speaker said each part of a recording.

It runs locally as a Python package with a command-line interface, and the repository includes a troubleshooting guide for CUDA and cuDNN setup. It suits people producing subtitles, interview transcripts or meeting notes who need precise timing and speaker labels rather than plain text.

Strengths

  • Word-level timestamps for subtitles and transcripts
  • Speaker diarization identifies who is talking
  • Runs on your own machine

Limitations

  • GPU setup with CUDA and cuDNN can take some troubleshooting
  • Command-line and Python only, no graphical interface
  • The research does not state the license

Details

Pricing
FreeFree to use from its GitHub repository.
License
Proprietary
Developer
Max Bain
Platforms
Command line
How it runs
Downloadable app
Best suited for
Transcribing interviews and meetings locally with timestamps and speaker labels
Last verified
Added
Provenance
Facts checked against the developer's own pages and store listings, 1 sources on file.

Alternatives to WhisperX

Compare all

Software that can replace WhisperX for an important use case, and what changes if you switch.

  • Whisper

    An open-source speech recognition model from OpenAI for transcribing and translating audio on your own machine.

    Whisper is OpenAI's MIT model that adds translation to English and language identification, but lacks WhisperX's word timestamps and diarization.

  • whisper.cpp

    C/C++ port of OpenAI's Whisper for fast offline speech-to-text on your own hardware.

    whisper.cpp is an MIT C/C++ port that runs offline on CPU or GPU without Python or CUDA setup, but has no speaker diarization.

  • noScribe

    Transcribes recorded audio to text on your own computer, with speaker separation.

    noScribe is GPL-3.0 with a desktop interface on Windows, macOS and Linux, and also separates speakers locally.

  • aTrain

    Transcribe recorded interviews on your own computer.

    aTrain is an AGPL-3.0 interview transcriber for Windows and Linux with speaker detection and no command-line setup required.

  • MacWhisper

    A Mac app that transcribes audio and video files to text on your own machine.

    FreemiumProprietarymacOS

    MacWhisper is a freemium, closed source Mac app with speaker recognition and a choice of Whisper, Qwen and Parakeet models.

  • Vibe

    Transcribe audio and video files on your own computer.

    Vibe is an MIT desktop app for Windows, macOS and Linux with batch jobs and several subtitle formats, plus optional cloud summarization.

  • Buzz

    Turn speech recordings into text with a desktop transcription app.

    Buzz is an MIT desktop app for Windows, macOS and Linux with microphone transcription and subtitle exports, without a terminal.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing