WhisperX
Whisper with fast forced alignment, accurate word-level timestamps, and multi-speaker diarization.
WhisperX
WhisperX extends OpenAI Whisper with phoneme-based forced alignment for word-level timestamps accurate to ±20ms, and integrates pyannote.audio for speaker diarization — letting you output speaker-labelled transcripts in one command.
Key Features
- Word-level timestamps: phoneme alignment via
wav2vec2gives far more accurate boundaries than Whisper's built-in timestamps - Speaker diarization: plug in a HuggingFace pyannote token to automatically label each segment by speaker
- Batched inference: chunked audio with
faster-whisperbackend for 70× real-time throughput on GPU - Language detection: automatic per-segment language ID for multilingual recordings
- SRT/VTT output: emit subtitle files directly from the CLI
- Minimal code change: drop-in replacement for
whisper.load_modelin existing pipelines
Quick Start
pip install whisperx
# Transcribe with word timestamps and speaker labels
whisperx audio.mp3 \
--model large-v3 \
--diarize \
--hf_token hf_xxx \
--output_format srt
import whisperx
model = whisperx.load_model("large-v3", device="cuda", compute_type="float16")
result = model.transcribe("audio.mp3", batch_size=16)
aligned = whisperx.align(result["segments"], ...)
npx ai-supply add whisperx-forced-alignment-diarization
Curated mirror of the open-source WhisperX (BSD-2-Clause). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/whisperx-forced-alignment-diarization/check). Click a policy:
Consume WhisperX programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/whisperx-forced-alignment-diarization
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/whisperx-forced-alignment-diarization/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add whisperx-forced-alignment-diarization
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/whisperx-forced-alignment-diarization/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "whisperx-forced-alignment-diarization" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.