Skip to content
ai-supply.store
DiscoverCategoriesLeaderboardsCommunityAgent APIFAQ
Sign inSign up free
catalog / Legal & Compliance / Apache OpenNLP — NLP Toolkit for Legal Documents
⬡PipelineLegal & ComplianceFree

Apache OpenNLP — NLP Toolkit for Legal Documents

Apache-licensed Java/Python NLP toolkit with tokenization, sentence detection, NER, POS tagging, and chunking — production-ready for legal document pipelines.

@ai-supply
Installs52k
⟳ upstream opennlp-3.0.0-M5 · updated 2d ago
↗ Source repository
← More Legal & ComplianceLegal & Compliance leaderboard →How we grade security →Source ↗
✓ Grade A · 100/100 · SafeSecurity assessment
✓No compromise signals16capabilities surfaced11of 20 OWASP controls clear
External endpoints declaredExternal endpoints declaredExternal endpoints declaredBroad capability surface
scanned 20h ago·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

Apache OpenNLP — NLP Toolkit for Legal Documents

Apache OpenNLP is a mature, production-grade NLP toolkit from the Apache Software Foundation. It provides a full suite of language processing components — tokenizer, sentence detector, part-of-speech tagger, named entity finder, chunker, parser, and coreference resolver — implemented as trainable maximum-entropy models. Widely used in legal document pipelines for court filing processing, regulatory text extraction, and contract analysis in enterprise Java environments.

Key Features

  • Full NLP pipeline: tokenize → sentence detect → POS tag → NER → parse
  • Trainable on domain-specific corpora (legal, medical, financial)
  • REST service via OpenNLP Sandbox for microservice deployments
  • Python bindings available via opennlp-python
  • Apache-2.0 — clear IP for commercial legal software

Quick Start

# Via opennlp Python wrapper
pip install opennlp

import opennlp
nlp = opennlp.OpenNLP("/path/to/models/")
sentences = nlp.sentence_detector("The plaintiff filed suit. The court denied relief.")
tokens = nlp.tokenizer(sentences[0])
pos_tags = nlp.pos_tagger(tokens)
print(list(zip(tokens, pos_tags)))
npx ai-supply add apache-opennlp

Curated mirror of the open-source Apache OpenNLP (Apache-2.0). Get it from the source.

Rating rank
#1
of 11 in Legal & Compliance
Install rank
#3
of 11 in Legal & Compliance
Security score
100/100 · A
safe
Security rank
#1
of 11 in Legal & Compliance
Installs
52k
cat avg 29k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the Legal & Compliance leaderboard →
✓ Security: Safe · 100100/100 · grade Ascanned 20h ago
✓ no compromise signals16 risk-surface · 4/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

What this capability can do · med confidence (static)
⚑ filesystem⚑ secrets
egress → www.apache.org, s.apache.org, opennlp.apache.org, img.shields.io, api.securityscorecards.dev, downloads.apache.org, bsky.app, the-asf.slack.com +32
38 steps⚑ uses secretswww.apache.orgs.apache.orgopennlp.apache.orgactions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1apache/infrastructure-actions/allowlist-check@maingithub.comactions/setup-java@03ad4de0992f5dab5e18fcb136590ce7c4a0ac95actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM05Improper Output Handlingmedium
Code that pipes model/user output into shell, eval, SQL or paths unsafely.
•Suspicious code patterns — dynamic code execution · apache-opennlp-a864230/opennlp-api/src/main/java/opennlp/tools/ml/model/MaxentModel.java (CWE-95)risk surface
⚠LLM06Excessive Agencymedium
Over-broad tool/permission surface or unrestricted egress.
•External endpoints declared — 3 distinct host(s) · apache-opennlp-a864230/.asf.yamlrisk surface
•External endpoints declared — 1 distinct host(s) · apache-opennlp-a864230/.github/CONTRIBUTING.mdrisk surface
•External endpoints declared — 2 distinct host(s) · apache-opennlp-a864230/.github/workflows/allowlist-check.ymlrisk surface
•Broad capability surface — 3 high-impact capability categories referenced — verify least-privilege · apache-opennlp-a864230/.github/workflows/publish-snapshots.yml (CWE-272)risk surface
•External endpoints declared — 14 distinct host(s) · apache-opennlp-a864230/NOTICErisk surface
•External endpoints declared — 12 distinct host(s) · apache-opennlp-a864230/README.mdrisk surface
•External endpoints declared — 4 distinct host(s) · apache-opennlp-a864230/opennlp-api/src/main/java/opennlp/tools/tokenize/WordpieceTokenizer.javarisk surface
•External endpoints declared — 7 distinct host(s) · apache-opennlp-a864230/opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/tokenize/lang/Factory.javarisk surface
•External endpoints declared — 19 distinct host(s) · apache-opennlp-a864230/opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/util/normalizer/UrlCharSequenceNormalizerCharacterizationTest.javarisk surface
•External endpoints declared — 5 distinct host(s) · apache-opennlp-a864230/opennlp-core/opennlp-runtime/src/test/resources/opennlp/tools/sentdetect/origin-training-data.txtrisk surface
•External endpoints declared — 6 distinct host(s) · apache-opennlp-a864230/opennlp-docs/src/docbkx/chunker.xmlrisk surface
•External endpoints declared — 16 distinct host(s) · apache-opennlp-a864230/opennlp-docs/src/docbkx/corpora.xmlrisk surface
⚠LLM10Unbounded Consumptionmedium
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
•Potentially unbounded loop — an infinite loop (while True / while(1) / for(;;)) may cause runaway consumption · apache-opennlp-a864230/opennlp-api/src/main/java/opennlp/tools/util/StringUtil.java (CWE-835)risk surface
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
✓LLM01Prompt InjectionPassed
✓LLM02Sensitive Information DisclosurePassed
✓LLM03Supply ChainPassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
✓LLM07System Prompt LeakagePassed
✓LLM08Vector and Embedding WeaknessesPassed
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
OWASP Machine Learning Security Top 10
⚠ML09Output Integritymedium
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
•Suspicious code patterns — dynamic code execution · apache-opennlp-a864230/opennlp-api/src/main/java/opennlp/tools/ml/model/MaxentModel.java (CWE-95)risk surface
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
✓ML02Data PoisoningPassed
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
✓ML05Model TheftPassed
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
✓ML06AI Supply ChainPassed
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (24) · hygiene / uncategorized
•Unrecognized file type — '.gitattributes' is not on the allowlist · apache-opennlp-a864230/.gitattributesrisk surface
•Unrecognized file type — '.gitignore' is not on the allowlist · apache-opennlp-a864230/.gitignorerisk surface
•Unrecognized file type — '.properties' is not on the allowlist · apache-opennlp-a864230/.mvn/wrapper/maven-wrapper.propertiesrisk surface
•Suspicious network references — suspicious TLD (2 URLs) · apache-opennlp-a864230/.mvn/wrapper/maven-wrapper.propertiesrisk surface
•Unrecognized file type — '.?' is not on the allowlist · apache-opennlp-a864230/LICENSErisk surface
•Disallowed file type — '.cmd' executables are not permitted · apache-opennlp-a864230/mvnw.cmd (CWE-434)risk surface
•Unrecognized file type — '.version' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-cli/src/main/resources/opennlp/tools/util/opennlp.versionrisk surface
•Unrecognized file type — '.sample' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/20newsgroup/sci.electronics/52794.samplerisk surface
•Unrecognized file type — '.conf' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/brat/brat-ann.confrisk surface
•Unrecognized file type — '.ann' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/brat/opennlp-1193.annrisk surface
•Unrecognized file type — '.conllu' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/conllu/de-ud-train-sample.conllurisk surface
•Unrecognized file type — '.hidden' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/leipzig/samples/.hiddenrisk surface
•Unrecognized file type — '.hdr' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/masc/fakeMASC.hdrrisk surface
•Unrecognized file type — '.sgm' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/muc/LDC2003T13.sgmrisk surface
•Unrecognized file type — '.sgml' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/muc/parsertest1.sgmlrisk surface
•Unrecognized file type — '.name' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/ontonotes/ontonotes-sample-01.namerisk surface
•Unrecognized file type — '.parse' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-formats/src/test/resources/opennlp/tools/formats/ontonotes/ontonotes-sample-02.parserisk surface
•Suspicious network references — suspicious TLD (1 URLs) · apache-opennlp-a864230/opennlp-core/opennlp-ml/opennlp-dl/README.mdrisk surface
•Unrecognized file type — '.dat' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-ml/opennlp-ml-maxent/src/test/resources/opennlp/tools/ml/maxent/football.datrisk surface
•Unrecognized file type — '.dict' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-runtime/src/test/resources/opennlp/tools/lemmatizer/smalldictionary.dictrisk surface
•Unrecognized file type — '.train' is not on the allowlist · apache-opennlp-a864230/opennlp-core/opennlp-runtime/src/test/resources/opennlp/tools/namefind/OnlyWithEntitiesWithTypes.trainrisk surface
•Disallowed file type — '.bat' executables are not permitted · apache-opennlp-a864230/opennlp-distr/src/main/bin/opennlp.bat (CWE-434)risk surface
•Disallowed file type — '.ps1' executables are not permitted · apache-opennlp-a864230/opennlp-distr/src/test/ps/test_opennlp.Tests.ps1 (CWE-434)risk surface
•Unrecognized file type — '.bats' is not on the allowlist · apache-opennlp-a864230/opennlp-distr/src/test/sh/test_opennlp.batsrisk surface
✔ verified source · pinned apache-opennlp-a864230 · changed since last scan (+5 pts)
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/apache-opennlp/check). Click a policy:

Consume Apache OpenNLP — NLP Toolkit for Legal Documents programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/apache-opennlp

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/apache-opennlp/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add apache-opennlp

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/apache-opennlp/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "apache-opennlp" })
OpenAPI spec →
vlatest
✓ Security: Safe · 1001mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Free, security-vetted AI capabilities — skills, MCPs, plugins, agents, datasets and more, each graded and freshness-tracked, and built for humans and agents alike.

api · v3.1status · all green
Contact
support@ai-supply.storesecurity@ai-supply.store
Catalog
  • Discover
  • Categories
  • Leaderboards
  • Benchmarks
  • Security
  • Scan a repo
Community
  • Community
  • FAQ
For agents
  • Quickstart (60s)
  • Authorize an agent
  • Agent API
  • OpenAPI spec
For builders
  • Publish
  • Dashboard
Account
  • Create account
  • Sign in
  • Settings
Legal
  • Terms
  • Publisher Agreement
  • Acceptable Use
  • Privacy