Skip to content
ai-supply.store
DiscoverCategoriesLeaderboardsCommunityAgent APIFAQ
Sign inSign up free
catalog / Cybersecurity / Presidio — PII Detection & Anonymization
⛨GuardrailCybersecurityFree

Presidio — PII Detection & Anonymization

Microsoft's open-source PII detection and anonymization engine supporting 50+ entity types across text, images, and structured data.

@ai-supply
Installs172k
⟳ upstream 2.2.364 · updated 4d ago
↗ Source repository
← More CybersecurityCybersecurity leaderboard →How we grade security →Source ↗
! Grade B · 88/100 · ReviewSecurity assessment
✓No compromise signals21capabilities surfaced1known CVE8of 20 OWASP controls clear
Broad capability surfaceBroad capability surfacePotentially unbounded loopVulnerable dependencies
scanned 20h ago·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

Presidio — PII Detection & Anonymization

Microsoft Presidio provides fast, contextual analysis and anonymization of personally identifiable information (PII) in text and images. It powers data-privacy compliance in LLM pipelines, ETL workflows, and document processing systems.

Key Features

  • 50+ built-in recognisers: names, emails, phone numbers, SSN, credit cards, IBANs, IP addresses, medical identifiers, and more
  • Custom recogniser support (regex, spaCy NER, stanza, transformer models)
  • Anonymization operators: redact, replace, hash, encrypt, mask, synthetic data substitution
  • Image redaction module (DICOM, PDF, raster)
  • REST API (presidio-analyzer + presidio-anonymizer as microservices)

Quick Start

from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine

analyzer = AnalyzerEngine()
results = analyzer.analyze(text="My phone is 212-555-1234", language="en")
anonymizer = AnonymizerEngine()
print(anonymizer.anonymize(text="My phone is 212-555-1234", analyzer_results=results))
# Output: My phone is <PHONE_NUMBER>
npx ai-supply add presidio-pii-anonymizer

Curated mirror of the open-source Presidio (MIT). Get it from the source.

Rating rank
#1
of 19 in Cybersecurity
Install rank
#3
of 19 in Cybersecurity
Security score
88/100 · B
review
Security rank
#6
of 19 in Cybersecurity
Installs
172k
cat avg 85k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the Cybersecurity leaderboard →
! Security: Review · 8888/100 · grade Bscanned 20h ago
✓ no compromise signals22 risk-surface · 7/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

Control card · high confidence (static)
framework: presidioframework: pytestframework: guardrails-aiframework: llm-guardcovers: piicovers: prompt-injectioncovers: secrets-leak
test_ssn_detectionvalidatetest_operator_is_non_reversibletest_operator_preserves_structuretest_when_valid_ssn_then_detect_with_correct_boundariestest_when_invalid_checksum_then_no_matchtest_when_context_missing_then_low_confidencetest_ssn_1test_case2predefinedcustomarraystringobjectnumberbooleanintegertype

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM03Supply Chainhigh
Vulnerable/compromised dependencies, models or archives in the artifact.
•Dependency manifest — 6 pip requirements declared · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/src/api/requirements.txtrisk surface
•Dependency manifest — 12 pip requirements declared · data-privacy-stack-presidio-f0dbdce/docs/samples/python/streamlit/requirements.txtrisk surface
•Dependency manifest — 4 pip requirements declared · data-privacy-stack-presidio-f0dbdce/e2e-tests/requirements.txtrisk surface
•Non-registry dependency source — 2 requirement(s) from git/URL/editable · data-privacy-stack-presidio-f0dbdce/e2e-tests/requirements.txt (CWE-829)risk surface
•Vulnerable dependencies — 7 known vulnerabilities in: transformers@4.57.6 (CWE-1395)known CVE · -12 pts
⚠LLM05Improper Output Handlinghigh
Code that pipes model/user output into shell, eval, SQL or paths unsafely.
•Suspicious code patterns — destructive rm -rf / · data-privacy-stack-presidio-f0dbdce/docs/samples/python/streamlit/Dockerfile (CWE-78)expected
•Suspicious code patterns — environment/secret exfiltration · data-privacy-stack-presidio-f0dbdce/e2e-tests/tests/test_package_e2e_integration_flows.py (CWE-200)expected
•Suspicious code patterns — dynamic code execution · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/tests/test_analyzer_engine.py (CWE-95)expected
•Suspicious code patterns — OS command execution; unsafe yaml.load · data-privacy-stack-presidio-f0dbdce/scripts/zensical_build.py (CWE-78)expected
⚠LLM06Excessive Agencyhigh
Over-broad tool/permission surface or unrestricted egress.
•External endpoints declared — 1 distinct host(s) · data-privacy-stack-presidio-f0dbdce/.github/PULL_REQUEST_TEMPLATE.mdexpected
•External endpoints declared — 3 distinct host(s) · data-privacy-stack-presidio-f0dbdce/.github/copilot-instructions.mdexpected
•Broad capability surface — 4 high-impact capability categories referenced — verify least-privilege · data-privacy-stack-presidio-f0dbdce/.github/workflows/ci.yml (CWE-272)risk surface
•External endpoints declared — 2 distinct host(s) · data-privacy-stack-presidio-f0dbdce/.github/workflows/codeql-analysis.ymlexpected
•Broad capability surface — 3 high-impact capability categories referenced — verify least-privilege · data-privacy-stack-presidio-f0dbdce/.github/workflows/release-docs.yml (CWE-272)risk surface
•External endpoints declared — 9 distinct host(s) · data-privacy-stack-presidio-f0dbdce/README.MDexpected
•External endpoints declared — 4 distinct host(s) · data-privacy-stack-presidio-f0dbdce/docs/analyzer/developing_recognizers.mdexpected
•External endpoints declared — 13 distinct host(s) · data-privacy-stack-presidio-f0dbdce/docs/community.mdexpected
•External endpoints declared — 5 distinct host(s) · data-privacy-stack-presidio-f0dbdce/docs/faq.mdexpected
•External endpoints declared — 6 distinct host(s) · data-privacy-stack-presidio-f0dbdce/docs/image-redactor/index.mdexpected
•Egress to a private/loopback host — 0.0.0.0 · data-privacy-stack-presidio-f0dbdce/docs/samples/docker/litellm.md (CWE-918)expected
•External endpoints declared — 8 distinct host(s) · data-privacy-stack-presidio-f0dbdce/docs/samples/python/langextract/index.mdexpected
•Egress to an anonymous-paste / tunnel / OOB endpoint — webhook.site · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/tests/test_url_recognizer.py (CWE-200)expected
⚠LLM07System Prompt Leakagemedium
Secrets, internal hosts or proprietary logic exposed in shipped prompts.
•Internal host / private infrastructure reference — shipped content references a private IP range or internal-only host · data-privacy-stack-presidio-f0dbdce/docs/samples/python/streamlit/azure_ai_language_wrapper.py (CWE-200)expected
⚠LLM10Unbounded Consumptionmedium
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
•Potentially unbounded loop — an infinite loop (while True / while(1) / for(;;)) may cause runaway consumption · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/docs/sample_for_presidio_pr/anonymization_toolkit_sample.md (CWE-835)risk surface
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
✓LLM01Prompt InjectionPassed
✓LLM02Sensitive Information DisclosurePassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
✓LLM08Vector and Embedding WeaknessesPassed
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
OWASP Machine Learning Security Top 10
⚠ML06AI Supply Chainhigh
Compromised PyPI/npm packages, typosquats, unsafe serialized models.
•Dependency manifest — 6 pip requirements declared · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/src/api/requirements.txtrisk surface
•Dependency manifest — 12 pip requirements declared · data-privacy-stack-presidio-f0dbdce/docs/samples/python/streamlit/requirements.txtrisk surface
•Dependency manifest — 4 pip requirements declared · data-privacy-stack-presidio-f0dbdce/e2e-tests/requirements.txtrisk surface
•Non-registry dependency source — 2 requirement(s) from git/URL/editable · data-privacy-stack-presidio-f0dbdce/e2e-tests/requirements.txt (CWE-829)risk surface
•Vulnerable dependencies — 7 known vulnerabilities in: transformers@4.57.6 (CWE-1395)known CVE · -12 pts
⚠ML09Output Integrityhigh
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
•Suspicious code patterns — destructive rm -rf / · data-privacy-stack-presidio-f0dbdce/docs/samples/python/streamlit/Dockerfile (CWE-78)expected
•Suspicious code patterns — environment/secret exfiltration · data-privacy-stack-presidio-f0dbdce/e2e-tests/tests/test_package_e2e_integration_flows.py (CWE-200)expected
•Suspicious code patterns — dynamic code execution · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/tests/test_analyzer_engine.py (CWE-95)expected
•Suspicious code patterns — OS command execution; unsafe yaml.load · data-privacy-stack-presidio-f0dbdce/scripts/zensical_build.py (CWE-78)expected
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
✓ML02Data PoisoningPassed
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
✓ML05Model TheftPassed
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (20) · hygiene / uncategorized
•Unrecognized file type — '.env' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/.envrisk surface
•Unrecognized file type — '.gitattributes' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/.gitattributesrisk surface
•Unrecognized file type — '.?' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/.github/CODEOWNERSrisk surface
•Unrecognized file type — '.gitignore' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/.gitignorerisk surface
•Unrecognized file type — '.helmignore' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/k8s/charts/presidio/.helmignorerisk surface
•Unrecognized file type — '.tpl' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/k8s/charts/presidio/templates/_helpers.tplrisk surface
•Suspicious network references — raw IP URL (3 URLs) · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/deployments/client/configmap.yamlexpected
•Unrecognized file type — '.mermaid' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/docs/.diagrams/architecture.mermaidrisk surface
•Unrecognized file type — '.bicep' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/infrastructure/main.biceprisk surface
•Unrecognized file type — '.sample' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/src/api/.env.samplerisk surface
•Unrecognized file type — '.dockerignore' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/docs/samples/deployments/openai-anonymaztion-and-deanonymaztion-best-practices/src/client_app/.dockerignorerisk surface
•Suspicious network references — raw IP URL (13 URLs) · data-privacy-stack-presidio-f0dbdce/docs/samples/docker/litellm.mdexpected
•Unrecognized file type — '.ini' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/e2e-tests/pytest.inirisk surface
•Unrecognized file type — '.dev' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/Dockerfile.devrisk surface
•Unrecognized file type — '.stanza' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/Dockerfile.stanzarisk surface
•Unrecognized file type — '.transformers' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/Dockerfile.transformersrisk surface
•Unrecognized file type — '.windows' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/Dockerfile.windowsrisk surface
•Unrecognized file type — '.j2' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/presidio_analyzer/conf/langextract_prompts/default_pii_phi_prompt.j2risk surface
•Unrecognized file type — '.typed' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-analyzer/presidio_analyzer/py.typedrisk surface
•Unrecognized file type — '.presidiocli' is not on the allowlist · data-privacy-stack-presidio-f0dbdce/presidio-cli/.presidioclirisk surface
✔ verified source · pinned data-privacy-stack-presidio-f0dbdce · changed since last scan (-12 pts)
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/presidio-pii-anonymizer/check). Click a policy:

Consume Presidio — PII Detection & Anonymization programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/presidio-pii-anonymizer

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/presidio-pii-anonymizer/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add presidio-pii-anonymizer

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/presidio-pii-anonymizer/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "presidio-pii-anonymizer" })
OpenAPI spec →
vlatest
! Security: Review · 881mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Free, security-vetted AI capabilities — skills, MCPs, plugins, agents, datasets and more, each graded and freshness-tracked, and built for humans and agents alike.

api · v3.1status · all green
Contact
support@ai-supply.storesecurity@ai-supply.store
Catalog
  • Discover
  • Categories
  • Leaderboards
  • Benchmarks
  • Security
  • Scan a repo
Community
  • Community
  • FAQ
For agents
  • Quickstart (60s)
  • Authorize an agent
  • Agent API
  • OpenAPI spec
For builders
  • Publish
  • Dashboard
Account
  • Create account
  • Sign in
  • Settings
Legal
  • Terms
  • Publisher Agreement
  • Acceptable Use
  • Privacy