Speech-to-speech (Symphony)
Need the lowest latency and the most natural turn-taking? Use Symphony, one model end to end.
- ~1 s to first agent audio
- Realtime tool calling
- Natural interruptions and turn-taking
Native speech-to-speech agents, plus the full STT → LLM → TTS pipeline when you need control at every step.
Talk to Symphony
Time to first agent audio
Pick one and try it here. Everything on this page runs on the same account you sign up for, and the voices carry between all four.
62 preset voices, each with its own origin and accent. Any of them can be asked for any of the 10 output languages below, and fidelity is strongest in the one they came from.
Output languages: English, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish.
Eight of 62. Every voice you clone or design joins them.
Browse the full librarySprag serves its own models on open-source infrastructure rather than reselling somebody else's endpoint. That is where the latency goes, and it is why our audio engineers can tune a voice for you when a stock one will not do.
Time to first agent audio by provider, lowest first: Sprag 1033 milliseconds (lowest), OpenAI 1458 milliseconds, Cartesia 1896 milliseconds, AssemblyAI 3107 milliseconds, Deepgram 3299 milliseconds.
1.18s p50 end-to-end turn latency on agentic_basic, n=30, against OpenAI at 1.76s and Cartesia at 2.24s. Speech in, speech out through a single model.
See the methodologyBoth run on the same API. Switch per use case, not per vendor.
Need the lowest latency and the most natural turn-taking? Use Symphony, one model end to end.
Need guardrails at the text layer, a specific LLM, or per-step logs? Use the cascade, with LLMs we serve and tune for speed, or your own.

Founded on Open-Source
−91.4%
reduction in job completion time arXiv↗
Omni-modal
Text, image, audio, and video inference in a single API layer.
Any architecture
Autoregressive, diffusion transformers, and non-autoregressive models in one engine.
OpenAI + ComfyUI
Fully OpenAI-compatible API server. Native ComfyUI support.
KV cache efficiency
State-of-the-art KV cache from vLLM core, applied to multimodal workloads.