Skip to content
ai-supply.store
ОбзорКатегорииРейтингиСообществоAgent APIFAQ
ВойтиБесплатная регистрация
catalog / Legal & Compliance / CaseHOLD — Legal Holdings Benchmark Dataset
▣DatasetLegal & ComplianceFree

CaseHOLD — Legal Holdings Benchmark Dataset

Stanford RegLab's 53K+ multiple-choice legal holdings dataset for training and evaluating legal NLP models.

@ai-supply
Установки13k
⟳ upstream main@daea97c · updated 3y ago
↗ Исходный репозиторий
← More Legal & ComplianceLegal & Compliance leaderboard →How we grade security →Source ↗
! Grade B · 75/100 · ReviewSecurity assessment
✓No compromise signals1known CVE10of 20 OWASP controls clear
Vulnerable dependencies
scanned 18d ago·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

CaseHOLD — Legal Holdings Benchmark Dataset

CaseHOLD is an expert-curated dataset of 53,000+ multiple-choice questions derived from US case law. Each question asks a model to identify the correct legal holding of a cited case from five candidate holdings, testing genuine legal reasoning rather than surface pattern matching.

Key features

  • 53,137 multiple-choice questions from federal court opinions (2010–2019)
  • Five candidate holdings per question — designed to require legal inference
  • Balanced across federal circuit courts (1st–11th plus DC)
  • Hosted on Hugging Face Datasets for one-line download
  • Reference fine-tuned models: custom-legalbert, legalbert-large available

Quick start

pip install datasets transformers
from datasets import load_dataset

ds = load_dataset("casehold/casehold", "all")
print(ds["train"][0])  # {citing_prompt, holding_0..4, label}

# Fine-tune your own legal LLM
from transformers import AutoTokenizer, AutoModelForMultipleChoice
tokenizer = AutoTokenizer.from_pretrained("nlpaueb/legal-bert-base-uncased")
npx ai-supply add casehold-legal-benchmark

Curated mirror of the open-source CaseHOLD (Apache-2.0). Get it from the source.

Rating rank
#1
of 11 in Legal & Compliance
Install rank
#7
of 11 in Legal & Compliance
Security score
75/100 · B
review
Security rank
#10
of 11 in Legal & Compliance
Installs
13k
cat avg 29k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the Legal & Compliance leaderboard →
! Security: Review · 7575/100 · grade Bscanned 18d ago
✓ no compromise signals1 risk-surface · 3/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

Data card · high confidence (static)
jsontxt
10 files

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM03Supply Chaincritical
Vulnerable/compromised dependencies, models or archives in the artifact.
•Vulnerable dependencies — 83 known vulnerabilities in: certifi@2020.12.5, filelock@3.0.12, idna@2.10.0, pyarrow@3.0.0, requests@2.25.1, torch@1.8.1, tqdm@4.49.0, urllib3@1.26.4 (CWE-1395)known CVE · -25 pts
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
◷LLM10Unbounded ConsumptionRuntime-enforced
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
✓LLM01Prompt InjectionPassed
✓LLM02Sensitive Information DisclosurePassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
✓LLM05Improper Output HandlingPassed
✓LLM06Excessive AgencyPassed
✓LLM07System Prompt LeakagePassed
✓LLM08Vector and Embedding WeaknessesPassed
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
OWASP Machine Learning Security Top 10
⚠ML06AI Supply Chaincritical
Compromised PyPI/npm packages, typosquats, unsafe serialized models.
•Vulnerable dependencies — 83 known vulnerabilities in: certifi@2020.12.5, filelock@3.0.12, idna@2.10.0, pyarrow@3.0.0, requests@2.25.1, torch@1.8.1, tqdm@4.49.0, urllib3@1.26.4 (CWE-1395)known CVE · -25 pts
⚠ML05Model Theftlow
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
•No license signal — no SPDX id or license keyword found · reglab-casehold-daea97c/LICENSErisk surface
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
◷ML09Output IntegrityRuntime-enforced
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
✓ML02Data PoisoningPassed
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (1) · hygiene / uncategorized
•Unrecognized file type — '.?' is not on the allowlist · reglab-casehold-daea97c/LICENSErisk surface
✔ verified source · pinned reglab-casehold-daea97c
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/casehold-legal-benchmark/check). Click a policy:

Consume CaseHOLD — Legal Holdings Benchmark Dataset programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/casehold-legal-benchmark

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/casehold-legal-benchmark/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add casehold-legal-benchmark

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/casehold-legal-benchmark/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "casehold-legal-benchmark" })
OpenAPI spec →
vlatest
! Security: Review · 751mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Бесплатные AI-возможности с проверкой безопасности — skills, MCP, плагины, агенты, датасеты и другое. У каждой своя оценка безопасности и контроль актуальности, и всё создано как для людей, так и для агентов.

api · v3.1status · all green
Контакты
support@ai-supply.storesecurity@ai-supply.store
Каталог
  • Обзор
  • Категории
  • Рейтинги
  • Бенчмарки
  • Безопасность
  • Scan a repo
Сообщество
  • Сообщество
  • FAQ
Для агентов
  • Быстрый старт (60s)
  • Авторизовать агента
  • Agent API
  • Спецификация OpenAPI
Для разработчиков
  • Опубликовать
  • Панель управления
Аккаунт
  • Создать аккаунт
  • Войти
  • Настройки
Правовые документы
  • Условия использования
  • Соглашение издателя
  • Правила допустимого использования
  • Конфиденциальность