PipelineDevOps & InfraFree

DVC

Git-like version control for ML datasets and pipelines — track experiments, reproduce results, and collaborate on data science projects.

Instalaciones61k
⟳ upstream 3.67.1 · updated 3mo ago
Repositorio fuente
Grade A · 100/100 · SafeSecurity assessment
No compromise signals19capabilities surfaced9of 20 OWASP controls clear
Broad capability surfacePotentially unbounded loopExternal endpoints declared · expectedExternal endpoints declared · expected
scanned 18d agoosv · gitleaks · opengrep · picklescan + heuristicsfull breakdown in the Security tab ↓

DVC — Data Version Control

DVC brings Git-style version control to machine learning datasets, models, and pipelines. Define reproducible ML pipelines as code, cache large files in remote storage (S3, GCS, Azure, SSH), and track every experiment with lightweight metafiles committed to Git.

Key features

  • Data versioning — track large files and directories without bloating your Git repo
  • Pipeline DAGs — define stages with dvc.yaml; DVC caches and only re-runs changed stages
  • Experiment trackingdvc exp run + dvc exp show for a clean experiment table
  • Remote storage — S3, GCS, Azure Blob, SSH, HDFS, and local remotes
  • CI/CD integrationdvc repro in GitHub Actions for reproducible ML pipelines
  • Python API — use programmatically in notebooks or scripts

Quick start

npx ai-supply add dvc-ml-pipeline-versioning

# Or install directly
pip install dvc

# Initialize in a Git repo
git init my-project && cd my-project
dvc init

# Track a dataset
dvc add data/train.csv
git add data/train.csv.dvc .gitignore
git commit -m "Track training data with DVC"

# Define a pipeline stage
dvc run -n train \
  -d data/train.csv -d src/train.py \
  -o model.pkl \
  python src/train.py

# Reproduce the pipeline
dvc repro

Curated mirror of the open-source DVC project (Apache-2.0). Install upstream from the repository.

More from @ai-supply

View profile →
Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
1.0M
Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
892k
Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
811k
Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
784k