Skip to content
ai-supply.store
탐색카테고리리더보드커뮤니티Agent APIFAQ
로그인무료 가입
catalog / DevOps & Infra / Opik (by Comet)
△EvalDevOps & InfraFree

Opik (by Comet)

Open-source LLM evaluation and observability platform — trace, test, and monitor LLM apps end-to-end.

@ai-supply
설치 수51k
⟳ upstream 2.1.32 · updated 11d ago
↗ 소스 저장소
← More DevOps & InfraDevOps & Infra leaderboard →How we grade security →Source ↗
✓ Grade A · 100/100 · SafeSecurity assessment
✓No compromise signals15capabilities surfaced6of 20 OWASP controls clear
Broad capability surfaceBroad capability surfaceExternal endpoints declared · expectedEmail addresses present · expected
scanned 9d ago · partial·osv · gitleaks · opengrep · picklescan + heuristics·full breakdown in the Security tab ↓

Opik — LLM Evaluation & Observability

Opik is an open-source end-to-end LLM evaluation platform by Comet. It provides distributed tracing for LLM applications, a rich evaluation framework for testing prompts and RAG pipelines, and production monitoring — all in a single self-hostable platform.

Key Features

  • LLM tracing: automatic instrumentation for LangChain, LlamaIndex, OpenAI, Anthropic
  • Evaluation datasets: version and manage evaluation datasets with golden answers
  • Automated scoring: built-in metrics (hallucination, answer relevance, context precision, BLEU)
  • LLM-as-judge: define custom scoring with any LLM
  • Online monitoring: track production metrics, catch regressions, set alerts
  • Self-hostable with Docker Compose or Kubernetes

Quick Start

import opik
from opik.integrations.openai import track_openai

opik.configure(use_local=True)  # self-hosted
client = track_openai(openai_client)

# All calls are now automatically traced
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What is RAG?"}]
)

Install via ai-supply

npx ai-supply add opik-llm-evaluation-platform

Curated mirror of the open-source Opik (Apache-2.0). Get it from the source.

Rating rank
#1
of 23 in DevOps & Infra
Install rank
#18
of 23 in DevOps & Infra
Security score
100/100 · A
safe
Security rank
#1
of 23 in DevOps & Infra
Installs
51k
cat avg 212k
This listing vs category average
Installs
this
cat avg
Security (of 100)
this
cat avg
Adoption trend
See the DevOps & Infra leaderboard →
✓ Security: Safe · 100100/100 · grade Ascanned 9d ago
✓ no compromise signals15 risk-surface · 8/20 OWASP controls flagged

Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.

Control card · high confidence (static)
framework: pytestframework: guardrails-aicovers: prompt-injectioncovers: secrets-leakcovers: pii
test_create_user__happyflowz.enumtest_validate_spantest_tracked_function__error_inside_inner_function__caught_in_top_level_spantest_optimization_lifecycle__happyflowtest_track__one_nested_function__happyflowtest_trace_creation__e2e__happyflowtest_sentiment_classificationmarkdowncheckboxestextarea

Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).

OWASP Top 10 for LLM Applications
⚠LLM01Prompt Injectionhigh
Adversarial instructions embedded in an artifact that hijack a downstream LLM.
•Prompt-injection phrasing — instruction-subversion language detected · opik/README.md (CWE-77)expected
⚠LLM02Sensitive Information Disclosurehigh
Secrets, credentials or PII shipped inside the artifact.
•Email addresses present — contains email-like strings · opik/.agents/commands/comet/create-and-run-tests.mdexpected
•Embedded credentials — found: Slack token · opik/.agents/docs/SLACK_MCP_SETUP.md (CWE-798)expected
•Phone number present — contains phone number-like pattern (E.164 or formatted) · opik/apps/opik-backend/Dockerfile (CWE-359)expected
⚠LLM05Improper Output Handlinghigh
Code that pipes model/user output into shell, eval, SQL or paths unsafely.
•Suspicious code patterns — destructive rm -rf / · opik/apps/opik-backend/Dockerfile (CWE-78)expected
⚠LLM07System Prompt Leakagehigh
Secrets, internal hosts or proprietary logic exposed in shipped prompts.
•Embedded credentials — found: Slack token · opik/.agents/docs/SLACK_MCP_SETUP.md (CWE-798)expected
⚠LLM06Excessive Agencymedium
Over-broad tool/permission surface or unrestricted egress.
•External endpoints declared — 1 distinct host(s) · opik/.agents/commands/comet/add-e2e-test.mdexpected
•External endpoints declared — 8 distinct host(s) · opik/.agents/commands/comet/generate-code-review-slack-command.mdexpected
•External endpoints declared — 5 distinct host(s) · opik/.agents/docs/SENTRY_MCP_SETUP.mdexpected
•External endpoints declared — 3 distinct host(s) · opik/.agents/docs/SLACK_MCP_SETUP.mdexpected
•External endpoints declared — 2 distinct host(s) · opik/.agents/skills/opik-external-integrations/references.mdexpected
•Broad capability surface — 3 high-impact capability categories referenced — verify least-privilege · opik/.agents/skills/python-sdk/testing.md (CWE-272)risk surface
•External endpoints declared — 4 distinct host(s) · opik/.github/workflows/lint_helm_chart.yamlexpected
•Broad capability surface — 4 high-impact capability categories referenced — verify least-privilege · opik/.github/workflows/typescript_sdk_integration_publish.yml (CWE-272)risk surface
•External endpoints declared — 10 distinct host(s) · opik/README.mdexpected
•External endpoints declared — 11 distinct host(s) · opik/apps/opik-backend/config.ymlexpected
⚠LLM08Vector and Embedding Weaknessesmedium
PII or plaintext source leakage in embedding/vector exports.
Embedding inversion/poisoning is largely runtime; static check covers PII in vector exports.
•Email addresses present — contains email-like strings · opik/.agents/commands/comet/create-and-run-tests.mdexpected
•Phone number present — contains phone number-like pattern (E.164 or formatted) · opik/apps/opik-backend/Dockerfile (CWE-359)expected
§LLM09MisinformationGovernance
Artifacts designed to produce false/deceptive output.
Detectable only by runtime behavioral evaluation; addressed via responsible-use attestation.
◷LLM10Unbounded ConsumptionRuntime-enforced
Unbounded loops/recursion causing DoS or runaway cost.
Enforced at runtime by the gateway (rate limits + spend caps + size caps); static check flags unbounded loops.
✓LLM03Supply ChainPassed
✓LLM04Data and Model PoisoningPassed
Backdoors/poisoning in training data or serialized models.
Behavioral poisoning needs model execution; static check covers unsafe serialization + dataset skew only.
OWASP Machine Learning Security Top 10
⚠ML02Data Poisoninghigh
Poisoned training datasets with triggers or anomalous distributions.
Static check covers trigger phrasing, PII and label skew; full poisoning detection is runtime.
•Email addresses present — contains email-like strings · opik/.agents/commands/comet/create-and-run-tests.mdexpected
•Prompt-injection phrasing — instruction-subversion language detected · opik/README.md (CWE-77)expected
•Phone number present — contains phone number-like pattern (E.164 or formatted) · opik/apps/opik-backend/Dockerfile (CWE-359)expected
⚠ML09Output Integrityhigh
Middleware tampering with model outputs in transit.
Gateway enforces TLS + response integrity; static check flags output-rewriting code.
•Suspicious code patterns — destructive rm -rf / · opik/apps/opik-backend/Dockerfile (CWE-78)expected
§ML01Input Manipulation (Adversarial)Governance
Models vulnerable to adversarial perturbations.
Requires runtime robustness evaluation; addressed via publisher robustness attestation.
§ML03Model InversionGovernance
Training data reconstructable from a model's outputs.
Runtime/evaluation property; addressed via model-card data-provenance + DP attestation.
§ML04Membership InferenceGovernance
Determining whether a record was in the training set.
Runtime/evaluation property; addressed via overfitting disclosure + DP attestation.
§ML08Model SkewingGovernance
Models trained on skewed data producing biased output.
Requires fairness evaluation; addressed via model-card bias/limitations disclosure.
✓ML05Model TheftPassed
Unlicensed re-distribution / license-incompatible derivatives.
Static check verifies license declaration; extraction throttling is runtime.
✓ML06AI Supply ChainPassed
✓ML07Transfer Learning AttackPassed
Backdoored base models / LoRA adapters propagating to derivatives.
Backdoor detection needs behavioral probing; static check covers unsafe serialization + provenance.
✓ML10Model Poisoning (Weights)Passed
Tampered model weight files; integrity must be verifiable.
Static check enforces safe formats + records a content hash for downstream verification.
Other findings (1) · hygiene / uncategorized
•Unrecognized file type — '.?' is not on the allowlist · opik/Makefilerisk surface
✔ verified source · pinned partial
Check against a policy

The same gate an agent runs before installing (POST /api/v1/trust/opik-llm-evaluation-platform/check). Click a policy:

Consume Opik (by Comet) programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/opik-llm-evaluation-platform

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/opik-llm-evaluation-platform/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add opik-llm-evaluation-platform

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/opik-llm-evaluation-platform/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "opik-llm-evaluation-platform" })
OpenAPI spec →
vlatest
✓ Security: Safe · 1001mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

무료로 제공하는 보안 검증 AI 역량 — skill, MCP, plugin, agent, 데이터셋을 비롯한 모든 항목에 보안 점수를 매기고 최신성을 추적하며, 사람과 agent 모두를 위해 만들었습니다.

api · v3.1status · all green
문의하기
support@ai-supply.storesecurity@ai-supply.store
카탈로그
  • 탐색
  • 카테고리
  • 리더보드
  • 벤치마크
  • 보안
  • Scan a repo
커뮤니티
  • 커뮤니티
  • FAQ
에이전트용
  • 빠른 시작 (60s)
  • 에이전트 승인
  • Agent API
  • OpenAPI 사양
빌더용
  • 게시
  • 대시보드
계정
  • 계정 만들기
  • 로그인
  • 설정
법적 정보
  • 이용약관
  • 게시자 계약
  • 이용 정책
  • 개인정보 처리방침