Skip to content
ai-supply.store
ExplorarCategoriasClassificaçõesComunidadeAgent APIFAQ
EntrarCadastre-se grátis
catalog / Data & ETL / Chonkie
⬡PipelineData & ETLFree

Chonkie

Lightweight, fast text-chunking library purpose-built for RAG ingestion, with token, sentence, semantic, and code-aware chunkers.

@ai-supply
Instalações31k
⟳ upstream v1.7.0 · updated 1mo ago
↗ Repositório fonte
← More Data & ETLData & ETL leaderboard →How we grade security →Source ↗
! Grade B · 75/100 · ReviewSecurity assessment
✓No compromise signals13capabilities surfaced1known CVE9of 20 OWASP controls clear
Suspicious code patternsSuspicious network referencesBroad capability surfaceSuspicious code patterns
scanned 1mo ago·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

Chonkie

Chunking is the unglamorous step that quietly decides RAG quality — split too coarsely and retrieval drags in noise, too finely and you shatter context. Chonkie is a no-nonsense library dedicated to doing exactly this one job well, with a tiny footprint and fast defaults instead of a heavyweight framework.

Key features

  • Multiple chunker strategies: token, word, sentence, recursive, semantic (embedding-similarity), and code-aware splitting
  • Built for speed and low memory so ingestion of large corpora stays cheap
  • Pluggable tokenizers and embedding backends; overlap and size fully configurable
  • Handy SemanticChunker groups semantically coherent sentences for cleaner retrieval units
  • Small dependency surface — installs light and drops into any ingestion pipeline

It slots in ahead of your embedding + vector-store step, turning raw documents into well-formed passages that downstream retrievers and rerankers can actually use.

Curated mirror of the open-source Chonkie (MIT). Get it from the source.

Rating rank
#1
of 24 in Data & ETL
Install rank
#20
of 24 in Data & ETL
Security score
75/100 · B
review
Security rank
#10
of 24 in Data & ETL
Installs
31k
cat avg 164k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the Data & ETL leaderboard →
! Security: Review · 7575/100 · grade Bscanned 1mo ago
✓ no compromise signals14 risk-surface · 6/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

What this capability can do · med confidence (static)
⚑ filesystem⚑ network⚑ secrets
egress → docs.chonkie.ai, discord.gg, docs.github.com, js.docs.chonkie.ai, huggingface.co, img.shields.io, pypi.org, codecov.io +28
51 steps⚑ uses secretsdocs.chonkie.aidiscord.ggdocs.github.comactions/checkout@v4pnpm/action-setup@v4actions/setup-node@v4actions/upload-pages-artifact@v3actions/deploy-pages@v4

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM03Supply Chaincritical
Vulnerable/compromised dependencies, models or archives in the artifact.
•Dependency manifest — 16 npm dependencies declared · feyninc-chonkie-0a6baea/docs/package.jsonrisk surface
•Vulnerable dependencies — 2 known vulnerabilities in: chromadb@1.5.9 (CWE-1395)known CVE · -25 pts
⚠LLM05Improper Output Handlinghigh
Code that pipes model/user output into shell, eval, SQL or paths unsafely.
•Suspicious code patterns — destructive rm -rf / · feyninc-chonkie-0a6baea/Dockerfile (CWE-78)risk surface
•Suspicious code patterns — environment/secret exfiltration · feyninc-chonkie-0a6baea/docs/scripts/fetch-releases.mjs (CWE-200)risk surface
•Suspicious code patterns — dynamic code execution · feyninc-chonkie-0a6baea/tests/chef/test_table_chef.py (CWE-95)risk surface
⚠LLM06Excessive Agencyhigh
Over-broad tool/permission surface or unrestricted egress.
•External endpoints declared — 2 distinct host(s) · feyninc-chonkie-0a6baea/.github/ISSUE_TEMPLATE/config.ymlexpected
•External endpoints declared — 1 distinct host(s) · feyninc-chonkie-0a6baea/.github/dependabot.ymlexpected
•External endpoints declared — 14 distinct host(s) · feyninc-chonkie-0a6baea/README.mdexpected
•External endpoints declared — 6 distinct host(s) · feyninc-chonkie-0a6baea/docs/.gitignoreexpected
•External endpoints declared — 3 distinct host(s) · feyninc-chonkie-0a6baea/docs/content/docs/chonkie/api/docker.mdxexpected
•Egress to a private/loopback host — 0.0.0.0 · feyninc-chonkie-0a6baea/docs/content/docs/chonkie/api/quickstart.mdx (CWE-918)expected
•Broad capability surface — 3 high-impact capability categories referenced — verify least-privilege · feyninc-chonkie-0a6baea/docs/content/docs/chonkie/experimental/chonkie-cli.mdx (CWE-272)risk surface
•External endpoints declared — 4 distinct host(s) · feyninc-chonkie-0a6baea/llms.txtexpected
⚠LLM10Unbounded Consumptionmedium
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
•Potentially unbounded loop — an infinite loop (while True / while(1) / for(;;)) may cause runaway consumption · feyninc-chonkie-0a6baea/src/chonkie/chunker/table.py (CWE-835)risk surface
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
✓LLM01Prompt InjectionPassed
✓LLM02Sensitive Information DisclosurePassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
✓LLM07System Prompt LeakagePassed
✓LLM08Vector and Embedding WeaknessesPassed
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
OWASP Machine Learning Security Top 10
⚠ML06AI Supply Chaincritical
Compromised PyPI/npm packages, typosquats, unsafe serialized models.
•Dependency manifest — 16 npm dependencies declared · feyninc-chonkie-0a6baea/docs/package.jsonrisk surface
•Vulnerable dependencies — 2 known vulnerabilities in: chromadb@1.5.9 (CWE-1395)known CVE · -25 pts
⚠ML09Output Integrityhigh
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
•Suspicious code patterns — destructive rm -rf / · feyninc-chonkie-0a6baea/Dockerfile (CWE-78)risk surface
•Suspicious code patterns — environment/secret exfiltration · feyninc-chonkie-0a6baea/docs/scripts/fetch-releases.mjs (CWE-200)risk surface
•Suspicious code patterns — dynamic code execution · feyninc-chonkie-0a6baea/tests/chef/test_table_chef.py (CWE-95)risk surface
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
✓ML02Data PoisoningPassed
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
✓ML05Model TheftPassed
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (8) · hygiene / uncategorized
•Unrecognized file type — '.gitignore' is not on the allowlist · feyninc-chonkie-0a6baea/.gitignorerisk surface
•Unrecognized file type — '.?' is not on the allowlist · feyninc-chonkie-0a6baea/Dockerfilerisk surface
•Unrecognized file type — '.mdx' is not on the allowlist · feyninc-chonkie-0a6baea/docs/content/docs/chonkie/api/docker.mdxrisk surface
•Suspicious network references — raw IP URL (6 URLs) · feyninc-chonkie-0a6baea/docs/content/docs/chonkie/api/quickstart.mdxrisk surface
•Unrecognized file type — '.mjs' is not on the allowlist · feyninc-chonkie-0a6baea/docs/next.config.mjsrisk surface
•Unrecognized file type — '.ini' is not on the allowlist · feyninc-chonkie-0a6baea/src/chonkie/api/alembic.inirisk surface
•Unrecognized file type — '.mako' is not on the allowlist · feyninc-chonkie-0a6baea/src/chonkie/api/migrations/script.py.makorisk surface
•Unrecognized file type — '.typed' is not on the allowlist · feyninc-chonkie-0a6baea/src/chonkie/py.typedrisk surface
✔ verified source · pinned feyninc-chonkie-0a6baea
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/chonkie-rag-chunking/check). Click a policy:

Consume Chonkie programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/chonkie-rag-chunking

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/chonkie-rag-chunking/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add chonkie-rag-chunking

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/chonkie-rag-chunking/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "chonkie-rag-chunking" })
OpenAPI spec →
vlatest
! Security: Review · 751mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Recursos de IA gratuitos e com segurança verificada — skills, MCPs, plugins, agents, datasets e muito mais, cada um com nota e acompanhamento de atualização, feitos tanto para pessoas quanto para agents.

api · v3.1status · all green
Contato
support@ai-supply.storesecurity@ai-supply.store
Catálogo
  • Explorar
  • Categorias
  • Classificações
  • Benchmarks
  • Segurança
  • Scan a repo
Comunidade
  • Comunidade
  • FAQ
Para agentes
  • Início rápido (60s)
  • Autorizar um agente
  • Agent API
  • Especificação OpenAPI
Para desenvolvedores
  • Publicar
  • Painel
Conta
  • Criar conta
  • Entrar
  • Configurações
Legal
  • Termos
  • Acordo de editor
  • Uso aceitável
  • Privacidade