Together AI is an Infrastructure & APIs company founded in 2022 and based in San Francisco, United States. It has raised $533.5M in total funding, most recently a Series B in 2025 at a $3.3B valuation.
| Date | Stage | Amount | Valuation | Lead investors |
|---|---|---|---|---|
| Feb 20, 2025 | Series B | $305M | $3.3B | General Catalyst, Prosperity7 Ventures |
| Mar 13, 2024 | Series A | $106M |
| $1.3B |
| Salesforce Ventures |
| Nov 29, 2023 | Series A | $102.5M | $500M | Kleiner Perkins |
| May 16, 2023 | Seed | $20M | — | Lux Capital |

Together AI raises $800M at an $8.3B valuation, up from $3.3B in early 2025.

Together AI closes an $800M Series C round to expand open-source model infrastructure and services.

Together AI presents nine papers at ICML 2026 spanning inference, training, and model optimization.

Together AI releases ParallelKernelBench, evaluating LLM performance on writing optimized multi-GPU CUDA kernels across 87 workloads.

Together AI optimizes MiniMax-M3 serving with sparse attention and paged decode to support 1M-token context and multimodal inference.
.png)
Together AI demonstrates the fastest speech-to-text stack on Artificial Analysis by optimizing the full system pipeline.
Serverless inference API providing on-demand access to 200+ open-source models spanning text generation, image generation (FLUX variants), code, embeddings, and audio modalities. Supports fine-tuning and batch processing with pay-per-use pricing, OpenAI-compatible interface, and zero data egress fees. Delivers up to 4x faster inference through proprietary optimization techniques including FlashAttention and custom NVIDIA Blackwell kernels, enabling developers to deploy production AI applications without managing infrastructure.
Self-service dedicated and instant GPU clusters built on NVIDIA H100, H200, B200, GB200, and GB300 hardware. Enables teams to provision clusters starting from single 8-GPU nodes with flexible 3-day minimum commitments. Includes Kubernetes or Slurm orchestration options, transparent pricing with no long-term contracts, and expert AI advisory support for model development. Optimized for training, fine-tuning, and inference at scale with up to 90% faster training operations powered by Together Kernel Collection.
Production-grade model customization supporting both LoRA and full weight fine-tuning for models up to 100B parameters. Enables teams to adapt open-source models on proprietary data for specialized tasks including legal, medical, and customer support applications. Features automatic evaluation during training, one-click deployment to serverless or dedicated endpoints, and transparent token-based pricing with no infrastructure management overhead required.
Secure, sandboxed execution environment for LLM-generated code and development workloads. Provides isolated compute for code interpretation and full development environments billed per vCPU hour. Enables safe execution of AI-assisted programming tasks at production scale, allowing developers to run code generated by language models without exposing systems to security risks.
1 patent on file, but none with both an extractable figure and an abstract on Google Patents yet.