Independent benchmarks
Speech model performance
Latency and accuracy across STT, TTS, and agentic models. Methodology in the open.
STT — 100 samples · LibriSpeech · run 2026-08-19
Latency vs accuracy
TTFT p50 on x, WER mean on y. Bottom-left is fastest and most accurate.
Latency vs cost
TTFT p50 on x, price per audio hour on y. Bottom-left is fastest and cheapest.
| Provider / Model | TTFT p50 | TTFT p95 | E2E p50 | RTF p50 | WER | CER | $/audio hr |
|---|---|---|---|---|---|---|---|
SpragSymphony | 149 | 308 | 264 | 0.043× | 0.9% | 0.3% | $0.09 |
SpragRhythm | 186 | 362 | 186 | 0.030× | 1.8% | 0.7% | $0.07 |
DeepgramNova 3 | 364 | 696 | 364 | 0.054× | 2.0% | 0.8% | $0.29 |
DeepgramNova 2 | 396 | 763 | 396 | 0.064× | 1.6% | 0.5% | $0.35 |
OpenAIGPT-4o Mini Transcribe | 565 | 1,110 | 565 | 0.089× | 1.3% | 0.5% | $0.18 |
ElevenLabsScribe v2 | 584 | 968 | 584 | 0.091× | 1.3% | 0.4% | $0.22 |
OpenAIGPT-4o Transcribe | 711 | 1,214 | 711 | 0.114× | 0.8% | 0.2% | $0.36 |
SonioxSTT Async v5 | 1,718 | 2,367 | 1,718 | 0.274× | 2.4% | 0.9% | $0.10 |
AssemblyAIUniversal-3.5 Pro | 2,238 | 3,813 | 2,238 | 0.358× | 1.0% | 0.2% | $0.21 |