RAGAS
Apache-2.0 RAG evaluation framework — faithfulness, answer relevancy, context recall, and more in one pip install.
RAGAS
RAGAS (Retrieval Augmented Generation Assessment) is an open-source framework for evaluating RAG pipelines end-to-end. It provides reference-free metrics that assess both the retrieval and generation stages without requiring ground-truth labels — making it practical for production monitoring as well as offline development.
Key features
- Reference-free metrics: faithfulness, answer relevancy, context precision, context recall, context entity recall
- End-to-end dataset evaluation: score entire test sets in one call
- LangChain and LlamaIndex native integrations
- LLM-as-judge architecture — configurable judge model
- CI/CD friendly — JSON output, thresholds, dataset tracking
- Apache-2.0 license
Quick start
pip install ragas
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_recall
from datasets import Dataset
data = {
"question": ["What year was Python created?"],
"answer": ["Python was created in 1991."],
"contexts": [["Python was created by Guido van Rossum and first released in 1991."]],
"ground_truth": ["1991"]
}
dataset = Dataset.from_dict(data)
result = evaluate(dataset, metrics=[faithfulness, answer_relevancy, context_recall])
print(result)
Install via ai-supply
npx ai-supply add ragas-rag-evaluation
Curated mirror of the open-source RAGAS (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/ragas-rag-evaluation/check). Click a policy:
Consume RAGAS programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/ragas-rag-evaluation
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/ragas-rag-evaluation/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add ragas-rag-evaluation
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/ragas-rag-evaluation/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "ragas-rag-evaluation" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.