Opik (by Comet)
Open-source LLM evaluation and observability platform — trace, test, and monitor LLM apps end-to-end.
Opik — LLM Evaluation & Observability
Opik is an open-source end-to-end LLM evaluation platform by Comet. It provides distributed tracing for LLM applications, a rich evaluation framework for testing prompts and RAG pipelines, and production monitoring — all in a single self-hostable platform.
Key Features
- LLM tracing: automatic instrumentation for LangChain, LlamaIndex, OpenAI, Anthropic
- Evaluation datasets: version and manage evaluation datasets with golden answers
- Automated scoring: built-in metrics (hallucination, answer relevance, context precision, BLEU)
- LLM-as-judge: define custom scoring with any LLM
- Online monitoring: track production metrics, catch regressions, set alerts
- Self-hostable with Docker Compose or Kubernetes
Quick Start
import opik
from opik.integrations.openai import track_openai
opik.configure(use_local=True) # self-hosted
client = track_openai(openai_client)
# All calls are now automatically traced
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "What is RAG?"}]
)
Install via ai-supply
npx ai-supply add opik-llm-evaluation-platform
Curated mirror of the open-source Opik (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/opik-llm-evaluation-platform/check). Click a policy:
Consume Opik (by Comet) programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/opik-llm-evaluation-platform
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/opik-llm-evaluation-platform/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add opik-llm-evaluation-platform
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/opik-llm-evaluation-platform/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "opik-llm-evaluation-platform" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.