What Modal does
Modal (Modal Labs, Inc.) is an AI infrastructure platform that provides “serverless” compute primitives for running machine learning and AI workloads—including low-latency inference, LLM fine-tuning, reinforcement learning, batch processing, and secure execution of untrusted/agent-generated code—without requiring customers to operate their own GPU clusters. Modal’s core developer workflow is code-first: users define application logic (and resource needs) and deploy it to Modal, which then schedules and scales execution on demand, with pricing described as usage-based (paying for compute time while code runs). Modal positions its differentiation around owning the underlying infrastructure layer (file/container runtime and scheduling) and around performance techniques aimed at rapidly scaling inference servers, including GPU- and CPU-side snapshot/restore and related system components. The company’s product suite is organized around primitives such as Inference, Sandboxes, Training, Notebooks, and Batch, built on a shared container platform and storage layer. In addition to self-serve deployment, Modal is also building toward production-grade inference “ownership” via endpoint-oriented workflows (for example, Modal Auto Endpoints, which are described as OpenAI API-compatible and controllable/observable). Modal also expands “agent execution environments” through Sandboxes (for isolated execution of untrusted code) and invests in scaling Sandbox concurrency to support agentic and RL use cases at very large volumes.