CaseHOLD — Legal Holdings Benchmark Dataset
Stanford RegLab's 53K+ multiple-choice legal holdings dataset for training and evaluating legal NLP models.
CaseHOLD — Legal Holdings Benchmark Dataset
CaseHOLD is an expert-curated dataset of 53,000+ multiple-choice questions derived from US case law. Each question asks a model to identify the correct legal holding of a cited case from five candidate holdings, testing genuine legal reasoning rather than surface pattern matching.
Key features
- 53,137 multiple-choice questions from federal court opinions (2010–2019)
- Five candidate holdings per question — designed to require legal inference
- Balanced across federal circuit courts (1st–11th plus DC)
- Hosted on Hugging Face Datasets for one-line download
- Reference fine-tuned models:
custom-legalbert,legalbert-largeavailable
Quick start
pip install datasets transformers
from datasets import load_dataset
ds = load_dataset("casehold/casehold", "all")
print(ds["train"][0]) # {citing_prompt, holding_0..4, label}
# Fine-tune your own legal LLM
from transformers import AutoTokenizer, AutoModelForMultipleChoice
tokenizer = AutoTokenizer.from_pretrained("nlpaueb/legal-bert-base-uncased")
npx ai-supply add casehold-legal-benchmark
Curated mirror of the open-source CaseHOLD (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/casehold-legal-benchmark/check). Click a policy:
Consume CaseHOLD — Legal Holdings Benchmark Dataset programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/casehold-legal-benchmark
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/casehold-legal-benchmark/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add casehold-legal-benchmark
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/casehold-legal-benchmark/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "casehold-legal-benchmark" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.