promptfoo
LLM eval + red-teaming framework — test prompts and models against custom assertions, compare providers, and catch regressions in CI.
promptfoo
promptfoo is a developer-first LLM evaluation and red-teaming framework. Write test cases with assertions, run them against any model or prompt variant, compare results side-by-side, and integrate automated evals into your CI pipeline — before bad outputs reach production.
Key features
- Declarative test cases — define expected outputs with string matchers, regex, JSON schema, semantic similarity, or LLM-as-judge
- Multi-provider comparison — benchmark the same prompts across OpenAI, Anthropic, Groq, Mistral, and local models simultaneously
- Red-teaming — automated adversarial probing for jailbreaks, prompt injection, PII leakage, and harmful content
- CI integration —
promptfoo evalexits non-zero on failures; ships a GitHub Actions example - Web UI — browser-based results viewer and diff tool
- Caching — reuse LLM responses to speed up iterative prompt development
Quick start
npx ai-supply add promptfoo-llm-eval
# Or install directly
npm install -g promptfoo
# Initialize a config
promptfoo init
# promptfooconfig.yaml
prompts:
- "Summarize the following: {{text}}"
providers:
- openai:gpt-4o
- anthropic:claude-opus-4-5
tests:
- vars:
text: "The quick brown fox jumps over the lazy dog."
assert:
- type: contains
value: fox
- type: llm-rubric
value: "The summary is concise and accurate."
promptfoo eval
promptfoo view
Curated mirror of the open-source promptfoo project (MIT). Install upstream from the repository.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/promptfoo-llm-eval/check). Click a policy:
Consume promptfoo programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/promptfoo-llm-eval
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/promptfoo-llm-eval/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add promptfoo-llm-eval
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/promptfoo-llm-eval/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "promptfoo-llm-eval" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.