Stanza — Stanford NLP Python Toolkit
Stanford NLP's Python library for tokenisation, sentence segmentation, NER, dependency parsing, and coreference across 70+ languages.
Stanza
Stanza is Stanford NLP's production-grade Python NLP toolkit covering the full linguistic analysis pipeline. With pre-trained neural models for 70+ human languages, it delivers accurate tokenisation, multi-word token expansion, POS tagging, lemmatisation, NER, and dependency parsing in one consistent API.
Key Features
- 70+ language models, including low-resource languages
- Full NLP pipeline: tokenise → MWT → POS → lemma → depparse → NER → coref
- BiLSTM neural architecture; UD-trained dependency parsers
- spaCy-compatible wrapper (
stanza.pipeline.core.Pipeline) - Named entity recognition: 18+ entity types
- Biomedical/clinical NLP models (PubMed, MIMIC-III trained)
- CoreNLP Java server bridge for Stanford CoreNLP features
Quick Start
import stanza
stanza.download("en") # download once
nlp = stanza.Pipeline(lang="en", processors="tokenize,mwt,pos,lemma,depparse,ner")
doc = nlp("Barack Obama was born in Hawaii. He was the 44th President.")
for sent in doc.sentences:
for word in sent.words:
print(f"{word.text:15s} POS={word.upos:6s} HEAD={sent.words[word.head-1].text if word.head > 0 else 'root'}")
for ent in sent.ents:
print(f" NER: {ent.text} [{ent.type}]")
Install via ai-supply
npx ai-supply add stanza-stanford-nlp-toolkit
Curated mirror of the open-source Stanza (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/stanza-stanford-nlp-toolkit/check). Click a policy:
Consume Stanza — Stanford NLP Python Toolkit programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/stanza-stanford-nlp-toolkit
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/stanza-stanford-nlp-toolkit/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add stanza-stanford-nlp-toolkit
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/stanza-stanford-nlp-toolkit/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "stanza-stanford-nlp-toolkit" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.