LLM Guard — Input/Output Security Toolkit
MIT-licensed security toolkit by ProtectAI that sanitizes LLM prompts and responses — blocking prompt injection, toxic content, PII leakage, and secrets.
LLM Guard — Input/Output Security Toolkit
LLM Guard is a comprehensive security layer for LLM-powered applications, providing both input (prompt) and output (response) scanners that can be dropped in-line with any LLM call. It is built and maintained by ProtectAI and is widely used in production AI pipelines.
Key Features
- Input scanners: prompt injection detector, ban-topics filter, ban-substrings, anonymize (PII), token limit enforcement, regex guardrails
- Output scanners: no-refusal detector, relevance check, JSON/code validation, sensitive-data redaction, factual consistency
- Synchronous + async APIs; OpenAI-compatible
- Integrates with LangChain, LlamaIndex, and bare-metal OpenAI clients
- Self-hosted — no data leaves your infrastructure
Quick Start
from llm_guard import scan_prompt, scan_output
from llm_guard.input_scanners import PromptInjection, Anonymize
from llm_guard.output_scanners import Sensitive
scanned_prompt, results = scan_prompt(
scanners=[PromptInjection(), Anonymize()],
prompt="Ignore previous instructions and...",
)
print(scanned_prompt, results)
npx ai-supply add llm-guard-input-output-security
Curated mirror of the open-source LLM Guard (MIT). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/llm-guard-input-output-security/check). Click a policy:
Consume LLM Guard — Input/Output Security Toolkit programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/llm-guard-input-output-security
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/llm-guard-input-output-security/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add llm-guard-input-output-security
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/llm-guard-input-output-security/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "llm-guard-input-output-security" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.