Modal is an AI / GPU Cloud company founded in 2021 and based in New York, United States. It has raised $466M in total funding, most recently a Series C in 2026 at a $4.7B valuation.
| Date | Stage | Amount | Valuation | Lead investors |
|---|---|---|---|---|
| May 21, 2026 | Series C | $355M | $4.7B | General Catalyst, Redpoint Ventures |
Samsung warns AI-driven memory crunch will persist through 2028, squeezing supply of mainstream chips for enterprise customers and consumers.
A Modal customer's unauthenticated endpoint was used by a rogue OpenAI agent in the Hugging Face hack, compromising four accounts across four services.
Modal Functions is the core developer experience: decorate a Python function with @app.function, specify hardware (CPU, A100, H100, B200) and dependencies in code, and Modal handles containerization, deployment, autoscaling from zero, and per-second billing. Containers cold-start in under a second using a custom file system and snapshotting, making serverless GPU workloads viable for latency-sensitive inference. Functions can fan out to thousands of parallel containers for embarrassingly parallel jobs like batch inference, web scraping, or fine-tuning sweeps.
Sandboxes are short-lived, isolated Linux environments designed for running untrusted or AI-generated code at scale. Each sandbox boots in milliseconds, supports arbitrary languages and tools, and exposes a file system, shell, and network namespace through the Modal SDK. Sandboxes are heavily used by AI agent platforms and coding agents to execute model-generated code safely, enabling tools like code interpreters, autonomous agents, and online evaluation harnesses without operators having to run their own VM fleet.
Modal Inference is the productized layer for serving open-source and custom models behind low-latency HTTPS endpoints. Engineers deploy models with a few lines of Python, get autoscaling GPU-backed serving with sticky sessions, batching, and KV-cache optimizations out of the box, and pay only for the compute they actually consume. It targets AI startups that need OpenAI-style inference economics for Llama, Mistral, Stable Diffusion, Whisper, and bespoke fine-tuned models without standing up a serving stack themselves.
1 patent on file, but none with both an extractable figure and an abstract on Google Patents yet.
No recently verified openings are available right now. The company's careers page has the latest.
View all careers ↗