ConnectorDevOps & InfraFree

BentoML

Build, ship, and scale AI services — unified framework from local development to production Kubernetes.

इंस्टॉल215k
⟳ upstream v1.4.39 · updated 2mo ago
सोर्स रिपॉज़िटरी
! Grade B · 75/100 · ReviewSecurity assessment
No compromise signals35capabilities surfaced1known CVE7of 20 OWASP controls clear
Potentially unbounded loopBroad capability surfacePossible obfuscationVulnerable dependencies
scanned 18d agoosv · gitleaks · opengrep · picklescan + heuristicsfull breakdown in the Security tab ↓

BentoML

BentoML is an open-source unified model serving framework that lets you build AI services from any ML framework and deploy them on any infrastructure. It handles the full lifecycle from packaging models into reproducible Bentos to autoscaling Kubernetes deployments with adaptive batching.

Key Features

  • Framework agnostic: PyTorch, TensorFlow, Keras, XGBoost, scikit-learn, LLMs, diffusion models
  • Adaptive micro-batching: automatically batch requests for optimal GPU throughput
  • Runners API: modular service composition with independent scaling
  • Bento packaging: reproducible bundles with model, code, dependencies, Dockerfile
  • BentoCloud integration: one-command deployment to managed inference infrastructure
  • Built-in OpenTelemetry, Prometheus metrics, and gRPC support

Quick Start

import bentoml

@bentoml.service
class SentimentAnalyzer:
    model = bentoml.models.get("sentiment:latest")

    @bentoml.api
    def classify(self, text: str) -> str:
        return self.model.predict([text])[0]
# Serve locally
bentoml serve sentiment_service:SentimentAnalyzer

# Build + containerize
bentoml build && bentoml containerize sentiment:latest

Install via ai-supply

npx ai-supply add bentoml-model-serving-framework

Curated mirror of the open-source BentoML (Apache-2.0). Get it from the source.

More from @ai-supply

View profile →
Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
1.0M
Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
892k
Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
811k
Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
784k