BentoML
Build, ship, and scale AI services — unified framework from local development to production Kubernetes.
BentoML
BentoML is an open-source unified model serving framework that lets you build AI services from any ML framework and deploy them on any infrastructure. It handles the full lifecycle from packaging models into reproducible Bentos to autoscaling Kubernetes deployments with adaptive batching.
Key Features
- Framework agnostic: PyTorch, TensorFlow, Keras, XGBoost, scikit-learn, LLMs, diffusion models
- Adaptive micro-batching: automatically batch requests for optimal GPU throughput
- Runners API: modular service composition with independent scaling
- Bento packaging: reproducible bundles with model, code, dependencies, Dockerfile
- BentoCloud integration: one-command deployment to managed inference infrastructure
- Built-in OpenTelemetry, Prometheus metrics, and gRPC support
Quick Start
import bentoml
@bentoml.service
class SentimentAnalyzer:
model = bentoml.models.get("sentiment:latest")
@bentoml.api
def classify(self, text: str) -> str:
return self.model.predict([text])[0]
# Serve locally
bentoml serve sentiment_service:SentimentAnalyzer
# Build + containerize
bentoml build && bentoml containerize sentiment:latest
Install via ai-supply
npx ai-supply add bentoml-model-serving-framework
Curated mirror of the open-source BentoML (Apache-2.0). Get it from the source.
Compromise signals — malicious or tampered code (leaked secrets, backdoors, a dropped executable) — reduce the score, and known dependency CVEs carry a bounded penalty (they warrant review but never QUARANTINE — update the dependency to clear). Other dangerous-by-capability traits are risk surface, expected for some capabilities. Every finding is mapped to its OWASP control below.
Findings mapped to the OWASP Top 10 for LLM Applications (2025) and the OWASP Machine Learning Security Top 10. Expand any flagged control for the exact findings — compromise reduces the score; expected/risk-surface do not, except a known CVE, which carries a small bounded penalty (high/critical → Review).
The same gate an agent runs before installing (POST /api/v1/trust/bentoml-model-serving-framework/check). Click a policy:
Consume BentoML programmatically. Authenticate with an API key or session — see Authorize an agent.
# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/bentoml-model-serving-framework
# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/bentoml-model-serving-framework/check \
-H "Content-Type: application/json" \
-d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'
# CLI
npx ai-supply add bentoml-model-serving-framework
# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/bentoml-model-serving-framework/install \
-H "Authorization: Bearer $AIM_KEY"
# MCP tool
install_listing({ "slug": "bentoml-model-serving-framework" })OpenAPI spec →Curated mirror — latest upstream source. See the repository for tagged releases.