Voice & Conversational Interfaces benchmarks
How performance is measured at the frontier · as of 2026-09-10 · 5 benchmarks
Accuracy
Word Error Rate
% WERTranscription errors across the Open ASR test suite — lower is better.
↓ lower is better·Open ASR Leaderboard
- #012.4%
- #022.6%
- #03
ibm-granite/granite-speech-5.0-470m-turboctc-nc 3.2% - #04
ibm-granite/granite-speech-5.0-470m-turboctc 3.4% - #053.4%
- #063.6%
- #073.7%
- #08
assemblyai/universal-3-5-pro 3.7% - #09
OpenMOSS-Team/MOSS-Transcribe-Diarize 3.8% - #103.9%
- #11
soundsgoodai/Zipformer-cr-ctc-transducer-XL-290M 4.2% - #124.3%
Speed
Throughput (RTFx)
× real-timeAudio-seconds transcribed per second of compute — higher is faster.
↑ higher is better·Open ASR Leaderboard
- #01
ibm-granite/granite-speech-5.0-470m-turboctc 12,946× - #02
ibm-granite/granite-speech-5.0-470m-turboctc-nc 12,762× - #03
nvidia/parakeet-tdt_ctc-110m 6,084× - #04
nvidia/parakeet-tdt-0.6b-v3 6,076× - #05
nvidia/parakeet-tdt-0.6b-v2 6,025× - #065,870×
- #07
nvidia/parakeet-rnnt-0.6b 5,431× - #085,023×
- #09
usefulsensors/moonshine-streaming-tiny 4,445× - #10
nvidia/parakeet-rnnt-1.1b 4,139× - #11
abr-ai/niagara-38m-batch.en 4,015× - #12
usefulsensors/moonshine-tiny 3,840×
Naturalness
- #011,578
- #021,574
- #031,562
- #041,558
- #051,557
- #061,557
- #071,547
- #081,544
- #091,539
- #101,536
- #111,528
- #12
Gradium TTS (pre-2026.08) 1,526
Economics
Transcription cost
$ / minutePublished speech-to-text API price per minute — lower is cheaper.
↓ lower is better·LiteLLM (open pricing)
- #01$0.0007
- #02$0.0020
- #03$0.0037
- #04$0.0043
- #05$0.0060
- #06$0.0060
TTS cost
$ / 1M charsPublished text-to-speech price per million characters — lower is cheaper.
↓ lower is better·LiteLLM (open pricing)
- #01$4
- #02$15
- #03$16
- #04$30
- #05$30
- #06$50
- #07$60
- #08$180