OpenAI Evals
MIT-licensed framework for evaluating LLMs and AI systems — build custom evals, run model comparisons, log results.
OpenAI Evals
OpenAI Evals is a framework for evaluating LLMs and LLM-powered systems, open-sourced by OpenAI under the MIT license. It provides a library of 1000+ existing evals alongside a structured way to build new ones — covering accuracy, safety, robustness, and task-specific performance. Evals can target any model via the OpenAI API or custom completion functions.
Key features
- 1,000+ ready-made evals — logic, coding, translation, factuality, safety
- Custom eval builder:
model_graded,basic,matcheval types - Model-graded evals use an LLM as judge for open-ended tasks
- YAML-based eval spec format — version-control your evaluations
- Multi-model comparison support for red-teaming and A/B testing
- MIT license — contribute or use commercially
Quick start
pip install openai evals
# Run a built-in eval
oaieval gpt-4o test-match
# Register and run a custom eval
cat > evals/registry/evals/my-eval.yaml << 'EOF'
my-eval:
id: my-eval.dev.v0
metrics: [accuracy]
my-eval.dev.v0:
class: evals.elsuite.basic.match:Match
args:
samples_jsonl: my_samples.jsonl
EOF
oaieval gpt-4o-mini my-eval
Install via ai-supply
npx ai-supply add openai-evals-framework
Curated mirror of the open-source OpenAI Evals (MIT). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/openai-evals-framework/check). Click a policy:
Consume OpenAI Evals programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/openai-evals-framework
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/openai-evals-framework/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add openai-evals-framework
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/openai-evals-framework/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "openai-evals-framework" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.