Skip to content

Independent benchmarks

Speech model performance

Latency and accuracy across STT, TTS, and agentic models. Methodology in the open.

STT — 100 samples · LibriSpeech · run 2026-08-19

Latency vs accuracy

TTFT p50 on x, WER mean on y. Bottom-left is fastest and most accurate.

0.0000.0070.0130.0200.02705011003150420052507TTFT p50 (ms)WER (mean)betterSoniox / STT Async v5Deepgram / Nova 3Sprag / RhythmDeepgram / Nova 2OpenAI / GPT-4o Mini TranscribeElevenLabs / Scribe v2AssemblyAI / Universal-3.5 ProSprag / SymphonyOpenAI / GPT-4o Transcribe

Latency vs cost

TTFT p50 on x, price per audio hour on y. Bottom-left is fastest and cheapest.

$0.000$0.101$0.202$0.302$0.40305011003150420052507TTFT p50 (ms)$ per audio hourbetterOpenAI / GPT-4o TranscribeDeepgram / Nova 2Deepgram / Nova 3ElevenLabs / Scribe v2AssemblyAI / Universal-3.5 ProOpenAI / GPT-4o Mini TranscribeSoniox / STT Async v5Sprag / SymphonySprag / Rhythm
Provider / ModelTTFT p50TTFT p95E2E p50RTF p50WERCER$/audio hr
SpragSymphony
1493082640.043×0.9%0.3%$0.09
SpragRhythm
1863621860.030×1.8%0.7%$0.07
DeepgramNova 3
3646963640.054×2.0%0.8%$0.29
DeepgramNova 2
3967633960.064×1.6%0.5%$0.35
OpenAIGPT-4o Mini Transcribe
5651,1105650.089×1.3%0.5%$0.18
ElevenLabsScribe v2
5849685840.091×1.3%0.4%$0.22
OpenAIGPT-4o Transcribe
7111,2147110.114×0.8%0.2%$0.36
SonioxSTT Async v5
1,7182,3671,7180.274×2.4%0.9%$0.10
AssemblyAIUniversal-3.5 Pro
2,2383,8132,2380.358×1.0%0.2%$0.21

Samples were drawn from the LibriSpeech dataset with each run evaluating the same utterances across STT models. Reference comparisons were drawn from the attached transcripts for each sample. Full methodology