DeepInfra is an AI / GPU Cloud company founded in 2022 and based in Palo Alto, United States. It has raised $133M in total funding, most recently a Series B in 2026.
| Date | Stage | Amount | Valuation | Lead investors |
|---|---|---|---|---|
| May 4, 2026 | Series B | $107M | — | 500 Global, Georges Harik |
A serverless, OpenAI-compatible API that gives developers on-demand access to a large catalog of open-source models spanning text generation, image generation, speech recognition, and embeddings. Users send requests without provisioning or managing any GPU hardware, paying per token or per request at prices DeepInfra positions well below hyperscaler rates. The service is built from DeepInfra's own GPU hardware up through the API layer, and is designed for production-scale workloads, processing nearly five trillion tokens per week across its customer base.
Dedicated GPU capacity for teams that need reserved hardware for high-volume or latency-sensitive inference rather than shared serverless endpoints. Customers can run their own or custom models on isolated GPUs while still using DeepInfra's managed serving stack, avoiding the overhead of building and operating their own inference infrastructure. The offering targets enterprises and scaleups running open-source and agent-driven AI workloads that want predictable performance, cost control, and freedom from proprietary-model vendor lock-in.
We don't have a live feed for this company's ATS. Their careers page has every open role.
View all careers ↗