DeepEval
Apache-2.0 LLM testing framework — pytest-style unit tests for RAG pipelines, chatbots, and LLM outputs.
DeepEval
DeepEval is an open-source LLM evaluation framework that brings software-testing discipline to AI systems. Written for Python and styled after pytest, it lets teams write unit tests for LLM outputs, RAG pipelines, and chatbot quality — with 14+ built-in metrics including G-Eval, RAGAS-style faithfulness, answer relevancy, hallucination detection, and toxicity.
Key features
- 14+ built-in metrics: G-Eval, answer relevancy, faithfulness, contextual precision/recall, hallucination, toxicity, bias, summarization
- pytest integration —
assertLLM outputs like code - RAG-aware: test retrieval quality, context relevancy, and generation faithfulness in one suite
- CI/CD ready with JUnit XML output and GitHub Actions support
- Apache-2.0 license — commercial use fully permitted
Quick start
pip install deepeval
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric
def test_answer_relevancy():
test_case = LLMTestCase(
input="What is RAG?",
actual_output="RAG stands for Retrieval-Augmented Generation.",
retrieval_context=["RAG is a technique that enhances LLM outputs with retrieved documents."]
)
metric = AnswerRelevancyMetric(threshold=0.7)
assert_test(test_case, [metric])
# Run your LLM test suite
deepevals test run test_llm.py
Install via ai-supply
npx ai-supply add deepeval-llm-testing
Curated mirror of the open-source DeepEval (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/deepeval-llm-testing/check). Click a policy:
Consume DeepEval programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/deepeval-llm-testing
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/deepeval-llm-testing/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add deepeval-llm-testing
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/deepeval-llm-testing/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "deepeval-llm-testing" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.