LM Evaluation Harness
EleutherAI's MIT-licensed unified benchmark suite — the de-facto standard for evaluating language models across 200+ tasks.
LM Evaluation Harness
LM Evaluation Harness is the canonical open-source framework for evaluating language models, developed by EleutherAI. It provides a unified interface to run 200+ benchmark tasks (MMLU, HellaSwag, ARC, TruthfulQA, GSM8K, and more) against any HuggingFace model, OpenAI API, or custom endpoint — making reproducible, comparable LLM evaluation easy.
Key features
- 200+ built-in tasks — MMLU, ARC, HellaSwag, TruthfulQA, GSM8K, HumanEval, WinoGrande, and more
- Plug-in architecture: evaluate local HF models, OpenAI API, vLLM, Anthropic, or custom backends
- Powers the Open LLM Leaderboard on HuggingFace
- Supports few-shot, zero-shot, and chain-of-thought modes
- MIT license — use in CI/CD pipelines, commercial workflows
Quick start
pip install lm-eval
# Evaluate Mistral-7B on MMLU (5-shot)
lm_eval --model hf \
--model_args pretrained=mistralai/Mistral-7B-v0.1 \
--tasks mmlu \
--num_fewshot 5 \
--device cuda:0
Python API
import lm_eval
results = lm_eval.simple_evaluate(
model="hf",
model_args="pretrained=microsoft/Phi-3-mini-4k-instruct",
tasks=["arc_easy", "hellaswag"],
num_fewshot=0,
)
print(results["results"])
Install via ai-supply
npx ai-supply add lm-evaluation-harness
Curated mirror of the open-source LM Evaluation Harness (MIT). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/lm-evaluation-harness/check). Click a policy:
Consume LM Evaluation Harness programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/lm-evaluation-harness
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/lm-evaluation-harness/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add lm-evaluation-harness
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/lm-evaluation-harness/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "lm-evaluation-harness" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.