Word Error Rate
% WERTranscription errors across the Open ASR test suite — lower is better.
- #01modulate/vfast3.1%
- #02HojoAI/Hojo-ASR-V13.1%
Voice AI’s enterprise land grab is being won by vertical depth, not horizontal scale—yet valuations still reward the latter.
Sierra’s Japan Breakthrough: 97% Resolution Resets the Bar for Enterprise Voice AI
ElevenLabs turns voice cloning into a music layer — the liquidity moat deepens
Deepgram plants its flag on the edge: Nova-3 lands on Snapdragon PCs
ElevenLabs’ Korea ambassador program: the voice layer’s liquidity moat tightens again
Sierra builds conversational AI agents for enterprise customer-experience workflows, replacing tier-1 support teams.
ElevenLabs builds real-time text-to-speech and voice-cloning models with support for 29 languages and ultra-low latency.
Parloa builds no-code conversational AI agents for European enterprise contact centers, specializing in voice and omnichannel automation.
Decagon builds AI customer-support agents that handle ticket resolution, live chat, and voice calls for mid-market and enterprise SaaS companies.
Sesame builds fast, high-fidelity voice models for conversational AI with a focus on multilingual support and low-latency inference.
LiveKit provides open-source infrastructure for real-time audio and video, powering voice-agent orchestration and WebRTC-based communication.
Cognigy builds enterprise conversational AI for customer service, supporting voice, chat, and omnichannel deployments.
PolyAI builds conversational AI for enterprise contact centers, specializing in restaurant reservations, hospitality, and logistics support.
Deepgram provides deep-learning-based automatic speech recognition and text-to-speech APIs for real-time transcription and voice synthesis.
Cartesia develops ultra-low-latency voice models including Sonic, achieving ~40ms time-to-first-audio for real-time conversational AI.
Vapi provides a developer platform for building and deploying voice AI agents with sub-second latency and flexible integrations.
Wispr Flow makes an AI voice-dictation app that turns natural speech into clear, formatted text in any application across Mac, Windows, iPhone, and Android.
Resemble AI provides voice cloning and custom text-to-speech APIs for gaming, entertainment, and call-center applications.
DeepL provides neural machine translation and launched DeepL Voice in 2024 for real-time speech-to-speech translation.
Dialpad provides a cloud business communications platform with built-in voice AI for real-time transcription, sentiment analysis, and agent coaching.
Cresta provides real-time AI coaching and automation for contact-center agents, analyzing voice and chat to improve sales and support outcomes.
Kore.ai provides an enterprise conversational AI platform for contact centers, covering voice, chat, and automation across 100+ channels.
Observe.AI provides voice analytics and agent-assist tools for contact centers, analyzing 100% of calls for compliance and coaching.
SoundHound AI builds a voice-native conversational AI platform powering in-car assistants, restaurant drive-thru and phone ordering, and enterprise agents through its Amelia platform.
AssemblyAI provides AI-powered speech-to-text APIs with sentiment analysis, speaker diarization, and real-time transcription features.
Voiceflow provides a collaborative platform for designing, prototyping, and deploying conversational AI agents across voice and chat.
Sanas builds real-time accent conversion technology for contact-center agents, translating non-native accents to standard dialects on live calls.
Yellow.ai provides a conversational AI platform for customer service and employee support with voice, chat, and automation features.
Hume AI builds empathic voice models that detect and generate emotional prosody in speech for more human-like conversational agents.
Bland AI offers a platform for deploying outbound and inbound phone-based conversational agents at scale.
KUDO offers a cloud-based platform for multilingual meetings and events with AI-assisted interpretation and real-time translation.
Synthflow offers a no-code platform for building voice AI agents, targeting SMBs and solo entrepreneurs without technical resources.
Speechmatics offers automatic speech recognition in 50+ languages with real-time transcription and on-premises deployment options.
WellSaid Labs creates synthetic voice avatars for corporate training, e-learning, and media production with studio-quality output.
Rasa builds open-source and enterprise conversational AI for on-premise voice and text assistants, popular in regulated industries.
Murf AI provides a studio-quality text-to-speech platform with 120+ AI voices for video production, e-learning, and marketing content.
Lindy builds AI personal assistants that handle email, scheduling, and voice calls for individuals and small teams.
Rime develops fast, expressive text-to-speech models optimized for conversational AI and live-agent experiences.
Retell AI provides APIs and SDKs for building conversational voice agents with sub-800ms latency and built-in telephony support.
Goodcall provides AI phone answering services for small businesses, serving 42,000+ barbershops, salons, and local service providers.
Phonic provides conversational voice surveys and qualitative research tools powered by AI transcription and sentiment analysis.
Dasha builds conversational AI for voice and text, offering a low-code SDK for developers building phone-based agents.
Arini builds AI phone agents for dental practices, automating appointment scheduling, insurance verification, and patient follow-up.
Fish Audio builds multilingual TTS and voice-cloning models with support for Chinese, English, and other Asian languages.
Soniox develops high-accuracy, low-latency speech recognition optimized for conversational AI and telephony applications.
Venture capital deployed · 2026 YTD
Sector market cap
Largest raise · trailing 12mo
Catalysts ahead · next 12mo
FTC filed suit in 2025 alleging deceptive marketing of business opportunities and false earnings claims. March 2026 settlement banned the owners from marketing business opportunities, imposed an $18M monetary judgment (largely suspended), and effectively put the company out of business. Platform was inactive by 2025 with key features broken and refunds stalled.
Audio-seconds transcribed per second of compute — higher is faster.
Crowd-voted naturalness of synthesized speech (TTS Arena).
Published speech-to-text API price per minute — lower is cheaper.
Published text-to-speech price per million characters — lower is cheaper.
Senior talent moves · 18mo
Open frontier bottlenecks
Companies tracked
OpenAI launches a platform for enterprises to deploy and manage AI agents across workflows.
Nothing on the calendar yet.
The agentic-voice wave blew up check sizes: the 2026 median is $250M versus $50M in 2024 and $65M in 2025. Sierra's $950M Series D at $15.8B, ElevenLabs' $500M, and Parloa's $350M top ten $100M+ rounds in the past twelve months, while Synthflow's $20M and Wispr Flow's $25M Series As mark the small end.
Capital rotated decisively up-stack: 2023–2024 was seed-and-A-heavy (18 of 32 rounds); since 2025 it is B-and-later (18 of 26), topped by five Series Ds since December — PolyAI, Decagon, Parloa, ElevenLabs, Sierra. Exits arrived too: NICE paid $955M for Cognigy and Meta took PlayAI off the board in July 2025.
Accel and Andreessen Horowitz lead the count with six apiece, but the 2026 mega-rounds belong to crossover money: Tiger Global and GV on Sierra, Sequoia on ElevenLabs and Sesame, Coatue and Index on Decagon. NVIDIA and Google's AI Futures Fund both debuted as voice leads in late 2025. Insight Partners hasn't led since 2022.