Skip to content
ai-supply.store
DécouvrirCatégoriesClassementsCommunautéAgent APIFAQ
Se connecterInscription gratuite
catalog / Cybersecurity / In-The-Wild Jailbreak Prompts
▣DatasetCybersecurityFree

In-The-Wild Jailbreak Prompts

CCS'24 dataset of 15,140 in-the-wild ChatGPT prompts including 1,405 real jailbreak prompts for training and benchmarking jailbreak detectors.

@ai-supply
Installations39k
⟳ upstream main@4f4031b · updated 1y ago
↗ Dépôt source
← More CybersecurityCybersecurity leaderboard →How we grade security →Source ↗
✓ Grade A · 100/100 · SafeSecurity assessment
✓No compromise signals2capabilities surfaced9of 20 OWASP controls clear
Prompt-injection phrasing · expectedSuspicious code patterns · expected
scanned 1mo ago·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

In-The-Wild Jailbreak Prompts — real-world LLM jailbreak dataset

This CCS 2024 dataset collects 15,140 ChatGPT prompts scraped from Reddit, Discord, prompt-sharing websites, and open datasets — including 1,405 verified jailbreak prompts gathered over roughly a year — the largest measurement study of in-the-wild jailbreaks at its release.

Key features

  • 1,405 real jailbreak prompts plus a large pool of benign prompts for contrastive evaluation
  • Sourced from four platforms with timestamps to study how jailbreaks evolve over time
  • A ready-made corpus for training or benchmarking prompt-injection and jailbreak detectors
  • Accompanied by analysis of prompt-sharing communities and attack effectiveness
  • Grounds red-team coverage in prompts that attackers actually used in the wild

Rather than synthetic attacks, this dataset gives defenders authentic adversarial inputs, making it a strong foundation for evaluating whether a guardrail catches the jailbreaks people really deploy.

Curated mirror of the open-source In-The-Wild Jailbreak Prompts (MIT). Get it from the source.

Rating rank
#1
of 19 in Cybersecurity
Install rank
#12
of 19 in Cybersecurity
Security score
100/100 · A
safe
Security rank
#1
of 19 in Cybersecurity
Installs
39k
cat avg 85k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the Cybersecurity leaderboard →
✓ Security: Safe · 100100/100 · grade Ascanned 1mo ago
✓ no compromise signals2 risk-surface · 5/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

Data card · high confidence (static)
csvlicense: detected
⚑ contains injection/jailbreak text — a risk surface when fed to an LLM (expected for a jailbreak dataset)
13 files

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM01Prompt Injectionhigh
Adversarial instructions embedded in an artifact that hijack a downstream LLM.
•Prompt-injection phrasing — instruction-subversion language detected · verazuo-jailbreak_llms-4f4031b/README.md (CWE-77)expected
⚠LLM05Improper Output Handlingmedium
Code that pipes model/user output into shell, eval, SQL or paths unsafely.
•Suspicious code patterns — dynamic code execution · verazuo-jailbreak_llms-4f4031b/code/ChatGLMEval/ChatGLMEval.py (CWE-95)expected
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
◷LLM10Unbounded ConsumptionRuntime-enforced
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
✓LLM02Sensitive Information DisclosurePassed
✓LLM03Supply ChainPassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
✓LLM06Excessive AgencyPassed
✓LLM07System Prompt LeakagePassed
✓LLM08Vector and Embedding WeaknessesPassed
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
OWASP Machine Learning Security Top 10
⚠ML02Data Poisoninghigh
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
•Prompt-injection phrasing — instruction-subversion language detected · verazuo-jailbreak_llms-4f4031b/README.md (CWE-77)expected
⚠ML09Output Integritymedium
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
•Suspicious code patterns — dynamic code execution · verazuo-jailbreak_llms-4f4031b/code/ChatGLMEval/ChatGLMEval.py (CWE-95)expected
⚠ML05Model Theftlow
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
•No license signal — no SPDX id or license keyword found · verazuo-jailbreak_llms-4f4031b/.gitignorerisk surface
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
✓ML06AI Supply ChainPassed
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (2) · hygiene / uncategorized
•Unrecognized file type — '.gitignore' is not on the allowlist · verazuo-jailbreak_llms-4f4031b/.gitignorerisk surface
•Unrecognized file type — '.?' is not on the allowlist · verazuo-jailbreak_llms-4f4031b/LICENSErisk surface
✔ verified source · pinned verazuo-jailbreak_llms-4f4031b
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/in-the-wild-jailbreak-prompts-dataset/check). Click a policy:

Consume In-The-Wild Jailbreak Prompts programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/in-the-wild-jailbreak-prompts-dataset

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/in-the-wild-jailbreak-prompts-dataset/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add in-the-wild-jailbreak-prompts-dataset

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/in-the-wild-jailbreak-prompts-dataset/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "in-the-wild-jailbreak-prompts-dataset" })
OpenAPI spec →
vlatest
✓ Security: Safe · 1001mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Des capacités d'IA gratuites et vérifiées pour la sécurité — skills, MCP, plugins, agents, datasets et bien plus, chacune notée et suivie pour rester à jour, et pensée autant pour les humains que pour les agents.

api · v3.1status · all green
Contact
support@ai-supply.storesecurity@ai-supply.store
Catalogue
  • Découvrir
  • Catégories
  • Classements
  • Benchmarks
  • Sécurité
  • Scan a repo
Communauté
  • Communauté
  • FAQ
Pour les agents
  • Démarrage rapide (60s)
  • Autoriser un agent
  • Agent API
  • Spécification OpenAPI
Pour les développeurs
  • Publier
  • Tableau de bord
Compte
  • Créer un compte
  • Se connecter
  • Paramètres
Mentions légales
  • Conditions
  • Accord éditeur
  • Utilisation acceptable
  • Confidentialité