← All companies

Inference / Company profile

Together AI

What Together AI does

Together AI is an AI infrastructure “neocloud” focused on delivering high-performance training and inference for open-weight (open source / open models) and custom models. The company positions its platform as an end-to-end stack that spans GPU capacity procurement, low-level kernel and systems optimizations, and developer-facing products for serverless and production-grade inference and fine-tuning. Together’s core users are AI application teams and organizations that want to run open models in production with strong latency/cost/performance economics, along with teams that need training/fine-tuning capabilities tied into the same inference ecosystem.

Product-wise, Together AI offers multiple deployment form factors for inference (including serverless endpoints and capacity/reservation-style options such as Provisioned Throughput) and also supports fine-tuning and training workflows for open and custom models. Together’s differentiation is strongly linked to its research-to-production pipeline in hardware-aware kernel optimization (e.g., FlashAttention-4 and kernel automation efforts), plus production feature work aimed at throughput, efficiency, and operational reliability.

Strategically, Together AI’s most recent capital raise (Series C announced July 1, 2026) is intended to accelerate expansion of its platform and compute capacity commitments, while continuing to push “open models” as a production default rather than a niche option. In parallel, the company has been shipping newer inference and infrastructure products (e.g., Provisioned Throughput and dedicated GPU cluster partnerships) and expanding model-lab partnerships (e.g., Moonshot AI for Kimi model releases) to improve developer experience and “day-zero” access to new open-weight models.

News

Company record

Jul 29, 2026 · Official · Together AI BlogTogether AI announces strategic partnership with Moonshot AI to natively serve Kimi models

Together announced a partnership with Moonshot AI to serve Kimi models (starting with Kimi K3) on Together’s inference products, with day-zero availability and a unified integration into Together’s model library.

Jul 29, 2026 · Official · Together AI BlogTogether AI announces strategic partnership with Moonshot AI to natively serve Kimi models (models landing across serverless and capacity options)

Together’s announcement states Kimi K3 is live on Together AI and that future open Moonshot models will be added going forward, with availability across multiple inference form factors.

Jul 20, 2026 · Official · Together AI BlogTogether AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

Together announced a partnership with Y Combinator to deliver a dedicated YC GPU cluster aimed at giving YC portfolio companies faster, dedicated access to compute for inference and training.

Jul 08, 2026 · Official · Together AI BlogOpen, convenient and predictable: Introducing Provisioned Throughput

Together introduced Provisioned Throughput, describing a reserved-inference form factor with token-based pricing and a 99% uptime SLA.

Jul 01, 2026 · Official · Together AI BlogTogether AI announces $800M Series C to accelerate the shift to open-source AI

Together AI announced an $800 million Series C, describing the move as acceleration toward open-source AI and expansion of its production platform, including compute capacity commitments.

Jun 10, 2026 · Official · Together AI BlogBuilding trust in enterprise AI: Together AI earns ISO 27001:2022 certification

Together announced it received ISO 27001:2022 certification, describing improved enterprise governance for production AI workloads.

Mar 05, 2026 · Official · Together AI BlogFlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

Together published a research/product-adjacent engineering post introducing FlashAttention-4 with a focus on maximizing overlap for asymmetric hardware scaling on NVIDIA Blackwell.

Mar 24, 2025 · Official · Together AI BlogIntroducing Together Chat: use DeepSeek R1 for free, hosted in North America

Together announced Together Chat, a consumer app using Together AI APIs for interacting with open models (including DeepSeek R1) hosted in North America.

Show 4 earlier updates
Date not disclosed · Official · Together AI BlogAnnouncing Together Inference Engine – the fastest inference available

Together introduced what it described as a new inference engine, positioning it as faster than existing alternatives for serverless deployments on comparable hardware.

Date not disclosed · Official · Together AI BlogAnnouncing Together Inference Engine 2.0 with new Turbo and Lite endpoints

Together announced Inference Engine 2.0 and described new Turbo/Lite serverless endpoint offerings and positioning relative to other providers.

Date not disclosed · Official · Together AI BlogTogether AI announces the availability of OpenAI open models on Together AI

Together announced OpenAI open models as available on Together AI, describing both serverless and dedicated endpoint options and compatibility with fine-tuning workflows.

Date not disclosed · Official · Together AI BlogTogether AI delivers fastest inference for the top open-source models

Together described performance work around Together Kernels and hardware-aware execution tuned for NVIDIA Blackwell architecture, including FlashAttention-4 integration and fused MoE kernels.

Source map · 2 recurring channels · 14 references

Still resolving: Newsroom · Official X

Funding

Latest disclosed valuation$8.3B
Tracked capital$1.3B5 sourced rounds
DateRoundRaisedValuationLead / investorsEvidence

What it builds

Together AI Platform

Together AI describes a full-stack platform for open-source AI that includes optimized training and model shaping along with large-scale production inference.

source ↗
Serverless Inference (Together inference endpoints)

Together provides serverless inference endpoints for running open and custom models in production via an API surface.

source ↗
Provisioned Throughput

Provisioned Throughput offers reserved inference capacity with token-based pricing and a 99% uptime SLA, intended to bridge the gap between best-effort serverless and reserved/dedicated capacity.

source ↗
Dedicated Inference / Dedicated Endpoints

Together provides dedicated inference options for production usage that require stronger control and capacity planning than serverless deployments.

source ↗
GPU Clusters

Together provides GPU cluster offerings designed for production workloads (including dedicated cluster partnerships).

source ↗
Fine-Tuning

Together supports fine-tuning workflows for open and custom models as part of its developer platform for training and deploying generative AI models.

source ↗

Key metrics

Milestones & partnerships

Jul 29, 2026
Moonshot AI partnership for native Kimi model serving

Together announced a partnership to provide day-zero access to Moonshot’s open-weight Kimi model releases on Together AI.

Moonshot AI
Jul 20, 2026
Partnership with Y Combinator for a dedicated YC GPU cluster

Together and Y Combinator announced a dedicated GPU cluster designed to help YC startups access compute for inference and training without long-term compute commitments.

Y Combinator
Jul 08, 2026
Launch of Provisioned Throughput

Together introduced Provisioned Throughput with reserved capacity semantics, token-based pricing, and a 99% uptime SLA.

Jul 01, 2026
Series C financing announced (compute scale-up)

Together AI announced an $800M Series C to accelerate the shift to open-source AI and expand compute capacity commitments.

Jun 10, 2026
ISO 27001:2022 certification

Together announced it received ISO 27001:2022 certification, describing enterprise-grade information security governance for production AI workloads.

Mar 05, 2026
FlashAttention-4 published for Blackwell-focused kernel co-design

Together published FlashAttention-4 describing algorithm and kernel pipelining co-design for asymmetric hardware scaling on NVIDIA Blackwell.

Mar 24, 2025
Launch of Together Chat consumer application

Together announced Together Chat, a consumer app using Together AI APIs and hosted in North America.

Date not disclosed
Together Inference Engine v1 introduced

Together introduced Together Inference Engine v1 and positioned it as providing faster inference for certain serverless deployment comparisons.

Date not disclosed
Together Inference Engine 2.0 introduced (Turbo/Lite endpoints)

Together announced Inference Engine 2.0 with Turbo and Lite endpoints and related deployment options through the Together API.

Date not disclosed
Model availability expansion for OpenAI open models

Together announced availability of OpenAI open models on Together AI, including serverless and dedicated endpoint options.

OpenAI