Category
Audio & Speech
TTS, STT, music, diarization.
6 listings
Every Audio & Speech listing on ai-supply is a free, open-source AI capability that we scan and grade A–D for security before it appears here — so you can adopt with confidence, not just a star count. This category spans 4 Pipelines, 1 Skill and 1 MCP server.
Security posture: 1 of 6 scanned rated safe · avg score 72/100.
⬡Pipeline
WhisperX
Whisper with fast forced alignment, accurate word-level timestamps, and multi-speaker diarization.
ai-supply
! B · 88⟳ 2mo agosecretsshellfs
↓ 224k
⬡Pipeline
NVIDIA NeMo — Scalable Speech & LLM Training Framework
NVIDIA's modular framework for training, fine-tuning, and deploying speech recognition, TTS, and large language models at scale.
ai-supply
! B · 75⟳ 3mo agosecretsshellnetwork
↓ 208k
⬡Pipeline
SpeechBrain
All-in-one conversational AI toolkit for ASR, speaker recognition, speech enhancement, and language identification.
ai-supply
! B · 75⟳ 3mo agosecretsshellnetwork
↓ 76k
◆Skill
librosa — Python Audio & Music Analysis Library
Python library for audio and music analysis: spectrograms, MFCCs, beat tracking, pitch detection, and feature extraction.
ai-supply
✓ A · 100⟳ 1y agonetworkfs
↓ 74k
⬡Pipeline
ESPnet
End-to-end speech processing toolkit covering ASR, TTS, speech translation, enhancement, and speaker diarisation.
ai-supply
! D · 16⟳ 3mo agosecretsshellnetwork
↓ 67k
◇MCP server
ElevenLabs MCP Server
Official ElevenLabs MCP server for text-to-speech, voice cloning, and audio transcription.
ai-supply
! B · 75⟳ 1mo agonetworkfs
↓ 6.4k