NVIDIA NeMo — Scalable Speech & LLM Training Framework
NVIDIA's modular framework for training, fine-tuning, and deploying speech recognition, TTS, and large language models at scale.
NVIDIA NeMo
NeMo is NVIDIA's end-to-end framework for developing and deploying state-of-the-art conversational AI, large language models, and speech models. It is built on PyTorch Lightning and supports distributed training across thousands of GPUs with tensor, pipeline, and data parallelism.
Key Features
- ASR: Conformer, Citrinet, FastConformer — SOTA word error rates
- TTS: FastPitch, HiFi-GAN, Mixer-TTS for natural speech synthesis
- NLP/LLM: GPT-style training, instruction tuning (SFT), RLHF, parameter-efficient fine-tuning
- Multimodal: vision-language alignment pipelines
- Collections: modular model collections for ASR, NLP, TTS, Vision
- Megatron-LM integration for ultra-large-scale training
- Deployment: NVIDIA TRT-LLM, Triton Inference Server export paths
Quick Start
import nemo.collections.asr as nemo_asr
# Load a pre-trained ASR model
asr_model = nemo_asr.models.EncDecCTCModelBPE.from_pretrained(
model_name="stt_en_conformer_ctc_large"
)
transcriptions = asr_model.transcribe(["podcast.wav"])
print(transcriptions[0])
Install via ai-supply
npx ai-supply add nemo-speech-and-llm-framework
Curated mirror of the open-source NeMo (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/nemo-speech-and-llm-framework/check). Click a policy:
Consume NVIDIA NeMo — Scalable Speech & LLM Training Framework programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/nemo-speech-and-llm-framework
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/nemo-speech-and-llm-framework/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add nemo-speech-and-llm-framework
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/nemo-speech-and-llm-framework/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "nemo-speech-and-llm-framework" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.