← All companies

Inference / Company profile

BentoML

What BentoML does

BentoML is an open-source and enterprise inference platform used by ML teams to package, deploy, and operate model inference in production. The company’s open-source core (the BentoML Python framework) focuses on turning Python-defined “Services” into deployable inference workloads, while its commercial platform (BentoCloud and the broader “Bento Inference Platform” marketed on bentoml.com) targets production needs such as deployment automation, observability, scaling behavior, and deployment control across cloud and on-prem environments.

BentoML’s technical approach centers on creating a standardized “Bento” artifact that bundles application code, dependencies, and model artifacts so teams can reproduce deployments across environments. The BentoML documentation describes “Bentos” as the standardized packaging format for AI/ML services and provides workflows for building, deploying, and containerizing those artifacts.

On the enterprise side, BentoCloud is positioned as an inference management platform and compute orchestration engine built on top of BentoML’s open-source serving framework. BentoCloud’s documentation and product messaging emphasize cloud deployment workflows, support for GPUs, and operating inference systems across development, testing, deployment, monitoring, and CI/CD.

As of February 2026, BentoML’s strategic position shifted materially: BentoML was acquired by (joined) Modular. BentoML and Modular both state that BentoML remains Apache 2.0 and that existing customers are supported without disruption while the teams integrate deeper into Modular’s stack.

News

Company record

Feb 10, 2026 · Official · BentoML BlogBentoML is joining Modular (product update and integration messaging)

BentoML published an update explaining that it joined Modular as part of a strategic product acquisition, emphasizing continued customer support and continued Apache 2.0 open-source status.

Feb 09, 2026 · Official · Modular ForumBentoML acquired by (joins) Modular

Modular announced it acquired BentoML and stated that BentoML remains open source under Apache 2.0 while bringing BentoML’s cloud deployment platform together with Modular’s optimization and inference stack.

Sep 11, 2025 · Official · BentoML Blogllm-optimizer and LLM Performance Explorer launch

BentoML announced llm-optimizer, an open-source tool for benchmarking and optimizing LLM inference with constraint-based filtering, and described a companion “LLM Performance Explorer” for browsing results.

Date not disclosed · Official · BentoML BlogBentoML 1.2 release (direct deployment workflow and UI/client updates)

BentoML announced BentoML 1.2, including a streamlined deployment workflow that allows developers to directly deploy to BentoCloud via a command-line workflow and additional BentoCloud UI/client improvements.

Date not disclosed · Official · BentoML BlogIntroducing BentoCloud

BentoML introduced BentoCloud as an inference management platform and discussed capabilities such as BYOC, multi-cloud deployment abstraction, and providing dedicated deployments with configurable inference behavior.

Source map · 2 recurring channels · 18 references

Still resolving: Newsroom · Official X

Funding

Latest disclosed valuationNot disclosed
Tracked capital$9M1 sourced round
DateRoundRaisedValuationLead / investorsEvidence
Jun 26, 2023Seed$9M

What it builds

BentoML Open-Source

BentoML is an open-source Python framework for building and serving AI/ML model inference workloads by defining Services and deploying them using Bento artifacts across local and production environments.

source ↗
BentoCloud

BentoCloud is BentoML’s inference management and compute orchestration platform for deploying and operating AI inference workloads in cloud environments, including a BYOC (Bring Your Own Cloud) option and production deployment workflows on top of the open-source serving framework.

source ↗
Bento Inference Platform

Bento Inference Platform is BentoML’s marketed enterprise offering that includes deployment and inference management capabilities (positioned around controlling and optimizing inference infrastructure), with the modular strategy now integrated into Modular.

source ↗

Milestones & partnerships