LexGLUE — Legal Language Understanding Benchmark
Multi-task benchmark for legal NLP with 7 datasets covering EURLEX classification, contract clause labeling, court judgement prediction, and more.
LexGLUE — Legal Language Understanding Benchmark
LexGLUE is the legal analogue of GLUE/SuperGLUE — a comprehensive benchmark spanning seven legal NLP datasets and tasks. It standardizes evaluation across EURLEX (EU legislation classification), ECHR (court judgement prediction), LEDGAR (contract provision classification), SCOTUS (US Supreme Court decision area), ECtHR (article violation prediction), ContractNLI (contract NLI), and CaseHOLD (legal holding identification).
Key Features
- 7 legal NLP tasks in a single evaluation harness
- Covers EU and US jurisdictions across legislation, contracts, and case law
- HuggingFace Datasets integration for easy loading
- Leaderboard tracking state-of-the-art Legal-BERT, RoBERTa-legal, and other models
- CC-BY-4.0 dataset license with public reproducibility
Quick Start
from datasets import load_dataset
# Load EURLEX classification task
dataset = load_dataset("coastalcph/lex_glue", "eurlex")
print(dataset["train"][0]["text"][:200])
print(dataset["train"][0]["labels"]) # Multi-label list
# Load ECHR court judgement prediction
scotus = load_dataset("coastalcph/lex_glue", "scotus")
print(scotus["test"][0])
npx ai-supply add lex-glue-legal-benchmark
Curated mirror of the open-source LexGLUE (CC-BY-4.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/lex-glue-legal-benchmark/check). Click a policy:
Consume LexGLUE — Legal Language Understanding Benchmark programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/lex-glue-legal-benchmark
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/lex-glue-legal-benchmark/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add lex-glue-legal-benchmark
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/lex-glue-legal-benchmark/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "lex-glue-legal-benchmark" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.