← All companies

Inference / Company profile

Baseten

What Baseten does

Baseten is an AI inference-infrastructure company that helps AI product teams deploy, operate, and scale production model workloads (including open-weight, fine-tuned, and custom models). The company positions its “Baseten Inference Stack” as the combination of inference runtimes and infrastructure optimizations, paired with workflows and tooling intended to reduce engineering time spent on deployment, autoscaling, observability, and reliability. Baseten’s platform supports multiple deployment models (Baseten-managed cloud, self-hosted, and hybrid) and targets latency-, throughput-, and uptime-sensitive applications where inference cost and operational risk are critical.

From a product standpoint, Baseten offers dedicated inference for high-scale workloads, “pre-optimized” Model APIs intended for rapid prototyping/evaluation and production-grade serving, and tooling for training/fine-tuning that can connect back to the same serving stack. Baseten also expands into distribution for closed-weight model labs via a dedicated “Baseten for Model Labs” offering. Strategic differentiation emphasized across Baseten’s materials is (1) production performance engineering across the inference pipeline (e.g., cache-aware routing, continuous batching/speculative techniques, optimized runtimes), (2) elastically scalable infrastructure including multi-cloud operation and enterprise controls, and (3) end-to-end developer workflows designed to manage model lifecycle concerns (rollouts, monitoring, versioning/observability) rather than leaving those to teams.

As of 2026-08-23, Baseten appears to be in a scale-up phase after multiple large financings within ~18 months (Series D through Series F). Baseten’s own recent materials cite sharp growth in revenue and inference volume and increasing emphasis on inference as a foundational layer across the AI stack, including post-training and “post-training + inference” workflows for specialized models.

News

Company record

Date not disclosed · Official · BasetenAnnouncing Baseten’s $1.5B Series F (and $13B valuation)

Baseten announced a $1.5B Series F, led and co-led by the investor group specified in the announcement, reporting $13B valuation and citing growth in revenue and inference volume.

Date not disclosed · Official · BasetenAnnouncing Baseten for Model Labs

Baseten launched “Baseten for Model Labs,” describing an offering to help closed-weight model labs distribute and monetize models using Baseten infrastructure plus distribution via the Baseten Model Library and a Frontier Gateway-style API layer.

Date not disclosed · Official · BasetenIntroducing the Baseten Frontier Gateway

Baseten introduced the Baseten Frontier Gateway as a managed inference gateway for AI/model labs, positioned as built on top of Baseten Dedicated Inference to let labs serve production-grade inference via a white-labeled API.

Date not disclosed · Official · BasetenHow we built the world’s fastest API for GLM-5.2 (performance and production-serving notes)

Baseten published a technical/performance-focused post describing how it built high-performance serving for GLM-5.2, including elements it attributes to its inference stack and production runtime optimizations.

Date not disclosed · Official · BasetenIntroducing the Baseten Delivery Network: Fast cold starts for big models

Baseten launched the Baseten Delivery Network (BDN), positioning it as an approach to eliminate cold-start bottlenecks for large models at scale by addressing multiple cold-start root causes.

Date not disclosed · Official · BasetenAnnouncing the acquihire of Inferless by Baseten

Baseten announced that the Inferless team is joining Baseten to accelerate innovation in inference infrastructure and to better support developers needing mission-critical AI applications with high performance, reliability, and cost efficiency.

Date not disclosed · Official · BasetenParsed + Baseten: Building Models That Touch Grass (Parsed joins Baseten)

Baseten announced that Parsed is joining Baseten, describing a combined approach to connect training systems with Baseten’s inference and post-training stack for iterative improvement from production data.

Date not disclosed · Official · BasetenAnnouncing Baseten’s $300M Series E (at $5B valuation)

Baseten announced a $300M Series E at a $5B valuation, led by IVP and CapitalG, with participation from NVIDIA and other investors listed in the announcement.

Show 4 earlier updates
Date not disclosed · Official · BasetenAnnouncing Baseten’s $75M Series C

Baseten announced raising $75M in Series C with co-leads IVP and Spark and additional investors specified in the post; it emphasizes R&D, global expansion, and team growth.

Date not disclosed · Official · BasetenAnnouncing Baseten’s $150M Series D

Baseten announced a $150M Series D led by BOND, with Jay Simons joining its Board; it lists participating investors and outlines a performance- and reliability-focused roadmap.

Date not disclosed · Official · BasetenAnnouncing our Series B ($40M)

Baseten announced an additional $40M, led by IVP and Spark, citing multi-cloud support, runtime integrations, and autoscaling/cold-start progress as part of its inference infrastructure narrative.

Date not disclosed · Official · BasetenAnnouncing our Series A

Baseten announced its Series A and describes a public beta launch intended to help data science and machine learning teams build full-stack applications powered by models without managing infrastructure complexity.

Source map · 4 recurring channels · 13 references

Funding

Latest disclosed valuation$2.2B
Tracked capital$2.2B9 sourced rounds
DateRoundRaisedValuationLead / investorsEvidence
UndatedSeed (co-led by Sarah Guo from Greylock Partners in 2019)
Not attributed
+1 investor

What it builds

Baseten Inference Stack

Baseten’s inference stack combines inference runtime optimizations and infrastructure mechanisms intended to make model serving fast, reliable, and cost-efficient. The company describes the stack as covering both “inference runtime” and underlying inference-optimized infrastructure, and it connects to Baseten’s multi-cloud deployment approach.

source ↗
Dedicated Inference

Baseten’s dedicated inference offering is positioned for high-scale workloads that need dedicated capacity and production-grade performance for open-source, custom, and fine-tuned models served at scale.

source ↗
Model APIs (including GLM-5.2 Fast)

Baseten provides production-first, pre-optimized Model APIs intended to ship/evaluate workloads quickly while delivering performance optimized for production inference. Baseten has also introduced a “Fast tier” on Model APIs with a dedicated variant for GLM-5.2.

source ↗
Baseten for Model Labs

Baseten for Model Labs is positioned as a distribution and monetization platform for closed-weight model labs, including a Frontier Gateway-style managed inference API layer and distribution via the Baseten Model Library.

source ↗
Baseten Delivery Network (BDN)

The Baseten Delivery Network is a system intended to reduce cold-start bottlenecks for large models at scale by addressing causes such as slow weight pulls from upstream storage, replica stampedes under load, and upstream availability dependencies.

source ↗

Milestones & partnerships

Jul 29, 2026
Launch of Baseten for Model Labs

Baseten launched Baseten for Model Labs to help closed-weight model labs distribute and monetize models, tying together a gateway/API approach and listing via the Model Library.

Jun 23, 2026
GLM-5.2 Fast performance engineering highlighted in production context

Baseten published a post describing how its inference stack optimizations contribute to high-performance serving for GLM-5.2 APIs.

Jun 22, 2026
Series F completed with $13B valuation claim

Baseten announced its $1.5B Series F, reporting a $13B valuation and attributing growth to inference demand and Baseten’s scale/performance focus.

Jun 12, 2026
Launch of Baseten Frontier Gateway

Baseten introduced the Frontier Gateway as a managed inference gateway for AI labs to serve Baseten-hosted models under their own domain, with production-grade inference infrastructure and governance.

Mar 19, 2026
Baseten Delivery Network (BDN) launched for faster cold starts

Baseten introduced the Baseten Delivery Network to reduce cold-start latency for large models by addressing weight-pull speed, replica stampedes, and upstream dependencies.

Mar 09, 2026
Inferless acquihire completed

Baseten announced that the Inferless team was joining Baseten to accelerate inference infrastructure innovation.

Inferless
Feb 25, 2026
Series C fundraise announced

Baseten announced a $75M Series C co-led by IVP and Spark and listed additional participating investors.

Feb 05, 2026
Parsed joins Baseten (training + inference unification narrative)

Baseten announced Parsed is joining Baseten with a stated intent to connect Parsed’s training systems and Baseten’s inference and training stack to create continuous improvement loops from production data.

Parsed
Feb 05, 2026
Series E raised at $5B valuation claim

Baseten announced a $300M Series E at a $5B valuation and listed the investor group.